How to use this guide
Send the questions in advance, ask for written answers, and score consistency between the written answer and what you hear on the call. Vagueness that survives a written request is the finding.
- Score each dimension out of five, and weight security, governance and integration highest — those are the ones that block deployment, not the ones that impress in a demo.
- Require evidence, not assertion. "We are HIPAA-compliant" is a claim. A BAA, a subprocessor list and a sample audit export are evidence.
- Ask the same questions of every vendor, including incumbents and internal build proposals.
- Record what was not answered. In our experience this column predicts implementation pain better than any score.
1. Security, HIPAA and business associate agreements
- Will you sign a BAA? Does it cover every subprocessor that can access PHI, including model providers?
- Publish your current subprocessor list. What notice do we get before it changes?
- Is provider-side training on our submitted data contractually disabled? Show the clause.
- How is data encrypted in transit and at rest, including vector stores, caches and logs?
- Which certifications do you hold, what is their scope, and when were they last assessed?
- Where does our data physically reside, and can we constrain residency?
Background on what each of these controls is doing: HIPAA-compliant generative AI. Our own answers: security, compliance and trust.
2. Data handling, retention and audit
- What exactly is retained — prompts, retrieved context, outputs, embeddings, logs — and for how long?
- Can you produce, for a named user and date range, every prompt, retrieval, model call and output?
- How do you handle a patient-driven deletion request that touches derived data such as embeddings?
- On termination, what is returned, what is destroyed, and on what timeline?
- Are audit logs immutable, and who inside your organisation can read our data?
3. Governance and human oversight
- Where is the human review step configured, and can it be removed by a non-engineer?
- How is an agent's data access scoped, and is it independent of the builder's own permissions?
- Can we see every agent in the estate, its owner, its data reach and its current model, in one view?
- What change control governs the move from build to production?
- How are model changes reviewed and recorded?
Our position: AI governance. Context on why scoping must be central: no-code AI agents.
4. Model flexibility and portability
- Which providers do you support today, and can we use our own internal or self-hosted models?
- Can we change the model behind a production agent without editing the agent? Demonstrate it.
- Where do prompts and tool definitions live — in the platform, or written against one provider's API?
- Do you provide normalised cost and token accounting across providers?
- If we leave, what do we take with us: prompts, tool definitions, evaluation sets, logs?
What a credible answer looks like: multi-model AI for healthcare.
5. Retrieval quality and citation accuracy
- Is retrieval filtered by the requesting user's permissions at query time, or is content filtered after generation?
- How are clinical documents chunked, and how do you keep a dose attached to its indication?
- Does every substantive claim in an answer carry a citation to a source passage with an effective date?
- How does the system behave when the answer is not in the approved sources?
- How do you evaluate retrieval quality separately from answer quality?
- How often is the index refreshed, and how are permission changes propagated?
Detail on each: enterprise RAG for hospitals and health systems.
6. Agent capability and autonomy
- What can an agent do without human review, and where is that boundary configured?
- Can agents write to the EHR? If so, under what controls — and should they?
- How are agent tool calls logged, and can a failing loop be bounded?
- Show us three agents running in production at a comparable organisation, and what they do.
- How many agents does your largest customer run, and who maintains them?
A grounded view of what is realistic: agentic AI for the healthcare sector.
7. Integration, identity and scalability
- How do you integrate with Epic or our EHR, and what is read versus written?
- Do you support SSO against our identity provider, with role mapping?
- What does the first integration take in our engineering hours, honestly?
- How does the tenth department differ from the first in effort?
- What is your uptime commitment, and what happens to an in-workflow agent when a model provider degrades?
8. Evidence, references and total cost
- Name production deployments at organisations comparable to ours, and let us speak to them unaccompanied.
- How long from contract to first agent in production at those sites?
- What did the customer's own team have to build?
- Distinguish licence cost from model consumption cost from our internal engineering cost.
- What has a customer asked you for that you could not deliver?
Ask the last one. A vendor that cannot name a limitation either does not know their product or is not being straight with you, and both are expensive to discover after signing.
Frequently asked questions
How should a health system evaluate an AI platform?
Across ten dimensions: security and HIPAA safeguards, BAAs and subprocessors, data handling and retention, governance and human oversight, model flexibility, retrieval and citation accuracy, agent autonomy boundaries, integration with the EHR and identity systems, evidence of comparable production deployments, and total cost including internal engineering effort. Send the questions in writing, ask every vendor the same set, and record what goes unanswered.
What are the biggest mistakes in healthcare AI procurement?
Evaluating on demo quality rather than governance; accepting "we are HIPAA-compliant" without asking which safeguards and what evidence; ignoring the internal engineering cost, which frequently exceeds the licence; and committing to a single model provider, which converts every future pricing or capability change into a migration project.
Should we build our own healthcare AI platform instead of buying?
Building the agent logic is not the hard part; building the compliance boundary, access scoping, audit, evaluation harness and change control is, and it has to exist before the first agent ships. A reasonable test is to scope that platform work honestly, in engineer-months, and compare it against a vendor that already has it in production — then decide whether it is differentiating work for your organisation.
Is this guide biased toward Inference Analytics?
The criteria are written to be answerable by any vendor, and several would be uncomfortable for us to answer badly — particularly those on limitations, portability and what a customer's own team must build. Use it against us alongside everyone else. Where we state our own position, it is linked and labelled as ours.