Clinical AI is moving from prediction to production: The hard part is everything around the model.

Key findings
Evidence from major medical institutions exposes a structural weakness in model-centered healthcare: the ability to generate a convincing answer is being asked to stand in for the ability to operate a reliable clinical system. The consequences extend from inaccurate records and unsafe recommendations to unpredictable costs and an expanding burden of human correction.
In April 2026, researchers at Mass General Brigham and Harvard Medical School published an evaluation of 21 large language models across 29 clinical cases. The models were strongest when the information had already been assembled and a final diagnosis was required. They were consistently weaker at the earlier work: identifying competing explanations and determining what needed to be investigated. The study exposed a gap between answering a completed case and navigating an uncertain one. JAMA Network
At Mount Sinai, researchers tested what happened when a fabricated medical detail was inserted into otherwise plausible clinical scenarios. The models repeatedly treated invented tests, signs and conditions as real. A targeted warning reduced those failures, but did not eliminate them. Nature
These weaknesses are being documented by institutions helping shape medical AI—not simply by organizations struggling to afford better technology.
The underlying problem is a mismatch between what a language model produces and what a healthcare institution must establish before acting.
A model produces an interpretation. Clinical work requires evidence, an appropriate decision, the authority to proceed and a reliable account of what happened afterward.
When one model is expected to interpret the record, fill information gaps, recommend an intervention, check its own reasoning and direct the next action, those responsibilities collapse into a single fallible process.
That is the foundational error in model-centered clinical AI: making the model responsible not only for proposing an answer, but for establishing that the answer is sufficiently grounded, permitted and complete.
The alternative is not to abandon language models. It is to stop treating them as the system.
Key findings
- Clinical records can become less reliable even when generated text looks useful. In a Stanford clinical pilot, feedback on 100 AI-generated summaries identified omissions in 25% and inaccuracies in 20%. The deployment included physician review, and no severe harm was reported. JAMA Network
- Prompt-based precautions leave significant residual risk. In Mount Sinai’s fabricated-detail stress test, a mitigation prompt reduced the average failure rate from 65.9% to 44.2%. Changing the sampling temperature did not materially resolve the problem. Nature
- A human reviewer is not an automatic safety barrier. A University of Michigan-led study found that biased AI predictions reduced clinicians’ diagnostic accuracy; image-based explanations did not significantly reverse the effect. DOI
- More elaborate model workflows can make spending harder to control. Anthropic reported approximately 15 times as many tokens for its multi-agent research system as for chat interactions. The relevant cost is the complete workflow, not the final answer. Anthropic
- A different architecture has already demonstrated value. At Amazon Pharmacy, a system combining language-model extraction with pharmacy logic and safety checks reduced direction-related near-miss events by 33% during an experimental production deployment. Nature
The medical record is not another prompt
A clinical record is not simply a large collection of text to be compressed. It is evidence on which subsequent decisions may depend.
Consider the difference between a symptom that was never discussed and a symptom the patient explicitly denied. Between a test that is pending and a result that is normal. Between a clinician considering a diagnosis and confirming it.
A summary can sound entirely reasonable while erasing those distinctions.
Stanford’s MedAgentBrief pilot illustrates both the usefulness of language models and the importance of keeping their output subordinate to a clinical process. The system generated draft hospital-course summaries for physicians to review and optionally use. It employed multiple processing stages, linked statements to source notes and was associated with reduced physician burnout. Yet clinicians still identified missing information and inaccuracies in the subset of summaries receiving feedback. JAMA Network
That is evidence for a bounded, reviewed application—not for treating generated summaries as automatically authoritative records.
The concern is what happens when an organization expands the role of that output. Imagine an unsupported detail entering a summary, being copied into a subsequent record and then becoming input to another AI-assisted decision. The next system may no longer see an uncertain interpretation. It sees something that looks like established clinical history.
This is a failure pathway that deployments should test explicitly.
The same issue appears when information changes. A recommendation based on yesterday’s laboratory result does not remain appropriate merely because its wording is persuasive today. A newly recorded medication, revised diagnosis or corrected result may change the decision.
The software therefore needs to preserve more than the answer. It needs to preserve what the answer depended on.
For healthcare buyers, the commercial consequence is direct: every unsupported or ambiguous statement creates potential verification work downstream. Someone must return to the original record, reconcile the discrepancy and determine whether it affected another decision.
Fluent text can move uncertainty out of sight without resolving it.
A dependable clinical system should do the opposite. It should make consequential uncertainty visible and route the work needed to resolve it. When essential information is unavailable, that should change what the system is allowed to recommend or do.
The absence of a fact must not become permission to invent the most plausible one.
A benchmark winner is not automatically ready for clinical responsibility
The Mass General Brigham evaluation matters because it examined clinical reasoning in stages rather than judging only the final answer. Its scoring required models to identify the appropriate options without introducing incorrect ones. Performance varied substantially by task, with differential diagnosis—the consideration of competing explanations—consistently weaker than final diagnosis. JAMA Network
The procurement implication is more important than another model leaderboard.
In a completed case, the evaluator has already selected the relevant information. In an evolving workflow, deciding which information is needed is part of the task.
A model that performs well after receiving the decisive result has not necessarily demonstrated that it would ask for that result, avoid premature conclusions or respond appropriately before the result exists.
This distinction applies beyond diagnosis. An impressive treatment explanation does not establish that prerequisites have been checked. A coherent care plan does not establish that it accounts for the latest patient information. A correct calculation does not establish that the right inputs were selected.
Our conclusion is that clinical evaluation must follow the responsibilities being assigned, not the sophistication of the output.
If the product is a summarization assistant, test faithful summarization. If it influences treatment decisions, test the relevant decision process. If it can initiate actions, test whether it respects permissions, detects changed conditions and handles incomplete execution.
The evidence for one responsibility cannot be borrowed to justify another.
A high score may justify further development or a carefully bounded deployment. It does not confer general clinical authority.
Safety instructions cannot be the safety architecture
Prompting is useful. The mistake is allowing it to become the primary mechanism enforcing consequential restrictions.
In Mount Sinai’s 2025 study, researchers created 300 physician-validated simulated cases, each containing a fabricated medical element. Across six models, instructions to use validated information and acknowledge uncertainty reduced the tendency to elaborate on the false material. Substantial failure remained. Nature
The distinction is simple: asking software not to invent evidence is different from checking whether its claims are supported.
The same distinction applies to permissions. “Obtain approval before proceeding” is an instruction. Preventing an unapproved action from occurring is a control.
This becomes particularly important when models read material from outside the application.
A December 2025 JAMA Network Open study used controlled prompt-injection simulations to manipulate medical advice. Attacks induced unsafe recommendations in 102 of 108 trials. The attack scenarios were constructed by researchers, but the result identifies a concrete vulnerability: content supplied to the model could redirect its behavior despite existing safeguards. JAMA Network
The UK National Cyber Security Centre reaches an architectural conclusion in its guidance on prompt injection. It recommends deterministic, non-LLM safeguards that constrain system actions, rather than relying only on keeping malicious content away from the model. National Cyber Security Centre
For healthcare, that means an imported note, retrieved document or external message must not be able to grant permissions, waive required checks or redefine who may act.
Giving a model access to a document does not make the document trustworthy. Adding a citation does not establish that the source supports the recommendation. Asking a second model to check the first does not establish independence or correctness without evidence that the checking process works.
The system should be designed so that a mistaken or manipulated interpretation has a limited route to consequences.
Models may interpret information. They should not be allowed to rewrite the conditions under which clinical actions are permitted.
Human review is not a universal repair service
“Clinician reviewed” is an important status. It is not a complete explanation of how risk was controlled.
In a randomized clinical-vignette study involving 457 clinicians, University of Michigan researchers found that systematically biased AI predictions reduced diagnostic accuracy by 11.3 percentage points. Providing image-based explanations alongside the biased predictions did not significantly mitigate the effect. Standard AI assistance improved performance, demonstrating that the quality of the assistance changed the result—not simply the presence of a clinician. DOI
This finding concerns predictive AI rather than an LLM, but it challenges a safeguard routinely proposed for both: the assumption that a human will reliably identify and reject the machine’s mistakes.
A meaningful review process needs to establish what the reviewer is deciding.
Are they checking the accuracy of a summary? Assessing whether a recommendation applies to this patient? Confirming that required evidence exists? Authorizing an action? Those are different jobs.
Combining them into one approval button does not make them one job.
A clinician should be able to see the source information, material conflicts, unresolved questions and the conditions supporting the proposed action. Otherwise, review requires reconstructing the entire case behind the output.
That is not efficient supervision. It is manual repair.
The commercial risk is that responsibility shifts to the clinical team while the work required to discharge that responsibility remains poorly specified. The vendor supplies a generated answer; the institution supplies the judgment, verification, reconciliation and follow-through necessary to use it.
The American Hospital Association has pushed back on this division of responsibility. In its December 2025 letter to the FDA, it argued that vendors and developers must remain responsible for the ongoing integrity of their tools and that evaluation requirements should account for hospitals’ resources. American Hospital Association
A better deployment does not remove human accountability. It gives accountable people a process they can actually supervise.
A generated plan is not a completed clinical workflow
Some of the most consequential failures occur after the model has produced a reasonable answer.
Duke’s Sepsis Watch implementation recognized that distinction. Its published account described the need to connect risk identification with assessment, coordination and completion of treatment steps. The deployed approach involved rapid-response nurses and treating clinicians, rather than treating the model’s signal as the intervention itself. JMIR Medical Informatics
This is a broader lesson for LLM-based agents.
An agent can describe a sequence of clinical tasks without establishing that any of them has been appropriately authorized or completed. Connecting it to software tools increases what it can attempt. It does not establish reliable execution.
Consider an illustrative workflow in which a request is sent to another system, but no clear confirmation returns. The action may have failed. It may also have succeeded while the confirmation was lost.
A model-centered process can be tempted to infer the most plausible outcome or repeat the request. A dependable workflow needs to reconcile what actually occurred before deciding what to do next.
The same applies when circumstances change between approval and execution. An approval based on an earlier record should not silently carry forward after information arrives that would materially alter the decision.
These are not requests for better prose or deeper reasoning. They are requirements for the surrounding software.
The system needs an enduring record of what is proposed, approved, attempted, completed and unresolved. It needs to identify who owns an exception and prevent an uncertain outcome from disappearing when a conversation moves on.
A model’s account of what it intended to do is not evidence that the clinical work happened.
That distinction should shape the product specification from the beginning—not be added after the first ambiguous incident.
Unreliability creates a second bill
The economics of model-centered healthcare extend beyond model prices.
The first bill pays for computation. The second pays for what happens when the output is incomplete, inconsistent, unsupported or difficult to verify.
Anthropic’s account of building its multi-agent research system provides a useful illustration of the computational side. It reported roughly four times the token consumption for agents compared with chat interactions, and roughly fifteen times for multi-agent systems. Its engineering team also encountered excessive delegation and agents searching indefinitely for nonexistent sources. Anthropic
These are first-party observations from a research system, not hospital spending estimates. The budgeting implication nevertheless follows directly: a price per token is not a price per completed workflow.
Long records, repeated retrieval, multiple models, revisions and retries can all change the amount of computation required. When the number of steps is left open-ended, the organization must budget for variable behavior rather than a fixed service.
Human correction creates a parallel expense.
At a volume of 100,000 tasks a month, thirty seconds of review per task represents approximately 833 staff-hours. Whether those hours are additional depends on the previous workflow, but they cannot be excluded from the comparison simply because the first draft appeared instantly.
There are three distinct costs to understand:
| Cost | What the institution must account for |
|---|---|
| Producing the output | Model usage, retrieval, tool calls, repeated attempts and infrastructure. |
| Making it usable | Review, corrections, record reconciliation, escalation and exception handling. |
| Keeping it dependable | Local evaluation, monitoring, change assessment, support and incident investigation. |
The economic failure occurs when the business case counts the first category while assuming the other two will be absorbed.
More checking may be justified. More expensive models may produce enough benefit to be worth using. But a deployment must demonstrate that the added effort improves reliability—not merely produces another answer.
The appropriate target is the full cost of completing the intended work to the required standard.
That also changes how spending limits should operate. A budget threshold must not silently abandon a clinical task. A runaway model loop should not be allowed to consume resources indefinitely. The workflow needs a defined point for stopping, escalating or handing the work back.
Uncertainty should produce a controlled exception, not an unlimited compute bill followed by manual cleanup.
Regulators are examining the same gap
The regulatory direction is not a blanket rejection of AI. It is a rejection of treating initial model performance as the entire safety case.
The FDA’s January 2025 draft guidance for AI-enabled device software proposes risk management across the product lifecycle. Its August 2025 final guidance on Predetermined Change Control Plans addresses how specified modifications can be developed, validated and implemented within an authorized plan. Together, they put attention on the ongoing behavior of the device, not just the output of a development-stage evaluation. U.S. Food and Drug Administration
NHS England’s June 2026 review of clinical safety standards identifies gaps involving AI governance, interactions between systems, cross-organizational collaboration and post-implementation monitoring. It distinguishes the responsibilities of health-IT manufacturers from those of organizations deploying and using the systems. NHS England
The commercial implications are substantial.
The buyer needs to know what happens when a model is updated, a workflow changes or a new population is introduced. The supplier needs to explain how problems will be detected, investigated and communicated. Responsibility cannot end at installation.
But governance must operate inside the workflow as well as around it.
A committee can approve a use case. It cannot personally inspect every future input. A policy can require clinical authorization. It cannot enforce that requirement unless the application does. A monitoring report can identify a problem afterward. It cannot, by itself, prevent the next inappropriate action.
This is why governance cannot be reduced to paperwork, and safety cannot be reduced to a prompt.
The institution needs controls that remain effective when the model is wrong.
The alternative is already visible in production
One of the clearest examples comes from medication instructions, where a seemingly small change in wording can change how a patient takes a drug.
Researchers at Amazon Pharmacy, with Stanford affiliation among the authors, developed MEDIC: a system that used a language model to extract the components of prescription directions, then applied pharmacy logic and safety checks to assemble the output. When essential information was missing or checks failed, the system could withhold a suggestion rather than complete the direction speculatively. Nature
During experimental integration into Amazon Pharmacy’s production workflow, the system reduced direction-related near-miss events by 33%. These were errors caught and corrected before reaching the patient. The work was conducted by researchers employed at Amazon Pharmacy. Nature
The important distinction is architectural. Language understanding was one component; it was not given unrestricted responsibility for producing and certifying the final instruction.
That is neither anti-AI nor a return to rigid automation everywhere.
It is a decision to use different mechanisms for different responsibilities: a model where interpretation is useful, explicit checks where requirements are known, and clinical expertise where judgment is necessary.
The contrast is not between an intelligent system and an unintelligent one. It is between intelligence operating inside a dependable process and intelligence being asked to substitute for that process.
Production-ready must mean more than “the model usually gets it right”
A serious clinical deployment should be evaluated against five questions.
What does the system know, and how does it know it? Recorded facts, inferred interpretations and unresolved information must remain distinguishable. A generated statement should not become authoritative merely because it has entered a polished interface.
What must be established before an action can proceed? Required information, appropriate permissions and clinically necessary reviews should affect execution. They should not exist only in documentation describing intended behavior.
What happens when circumstances change? A material update should trigger the appropriate reassessment. A previous recommendation or approval must not be treated as permanently valid.
How does the system establish what actually happened? Requested, authorized, performed and verified are different statuses. Uncertain outcomes need reconciliation, not a plausible narrative.
What does failure cost, and who handles it? The buyer should understand the computational limits, review burden, exception process and division of responsibility before deployment expands.
These are proposed evaluation criteria, not a promise that a particular architecture eliminates clinical risk. Incorrect rules remain incorrect when enforced consistently. Clinical validity, professional accountability and ongoing evaluation still matter.
The change is where responsibility sits.
The model may suggest an action. It should not be the sole mechanism establishing that the evidence is sufficient, the action is permitted and the task has been completed.
A deployment that cannot explain those boundaries is not ready merely because its demonstration is impressive.
Healthcare needs system-centered AI
The evidence does not point to a simple choice between rejecting language models and trusting them with the clinical process.
It points to a more demanding standard.
Language models can contribute useful interpretation, extraction and drafting. But every increase in their responsibility needs corresponding evidence about the system’s behavior: how it handles unsupported information, how it constrains action, how it responds to change and how it recovers from failure.
That standard applies equally to an incumbent EHR vendor, a foundation-model company, a health-system innovation team and a new infrastructure provider.
No institution’s reputation makes the model infallible. No product label makes a generated answer authoritative. No disclaimer transfers away the work required to use it responsibly.
For buyers, the decisive question should therefore move from:
How impressive is the model?
to:
What clinical responsibilities does this system assume—and what continues to protect the patient when its model is wrong?
That is the question around which healthcare AI should be built.
Not a model at the center with clinicians compensating for its limits, but a clinical system with explicit responsibilities, dependable controls and models operating within them.
The foundational mistake is not putting AI into healthcare. It is putting the model in charge of responsibilities that belong to the system.
Research methodology
This article is Tensor’s analysis of publicly available clinical deployment studies, controlled evaluations, institutional publications, engineering reports and official guidance reviewed through October 1, 2026. The supplied deployment and regulatory dossiers informed the research agenda; their claims were checked and supplemented using original sources. Their framing of implementation costs and governance burdens was treated as a starting point for investigation, not as independently validated market data. Clinical AI Deployment Costs
Prospective deployments, clinical-vignette studies and adversarial stress tests are identified in the text. Their percentages describe the tested systems and conditions, not universal error rates or rates of patient injury. The institution names establish where research and deployment took place; they do not constitute a ranking of failure rates.
Cost examples are labeled through their context, and model-provider engineering figures are not presented as healthcare-wide averages. FDA draft recommendations are distinguished from final guidance. The article’s architectural and procurement conclusions are editorial analysis; the cited sources do not evaluate or endorse AIVA OS.
