AI beating doctors on paper case studies doesn't show it helps real patients in real clinics, so what would?
What evidence would prove medical AI actually works in clinics?
This explores what kind of study would convince you that a medical AI helps real patients in real clinics, as opposed to scoring well on test cases. It also asks which kinds of evidence in the collection come closest.
This explores what would count as real proof that medical AI works in clinics, not just on paper. The collection's short answer is that there is a ladder of evidence and most of the impressive headlines sit on its lower rungs. On the bottom rung are written case studies and simulated patients. Here the results are striking. One LLM beat baselines drawn from hundreds of physicians on diagnosis, triage and management in vignettes and an emergency room study Can language models reason better than physicians at diagnosis?. Google's AMIE beat primary care doctors on 28 of 32 specialist-rated measures in text-based simulated consultations Can an AI system diagnose better than primary care doctors?. But a simulated patient can't get worse, and a vignette has no waiting room. Its edge was in reasoning from the information gathered, not in drawing that information out of the patient. That second skill is the one real patients test hardest.
The next rung asks whether AI makes clinicians better, not whether it beats them. With LLM access, clinicians got the right diagnosis into their top-10 list more often on hard NEJM cases: 51.7% versus 36.1% without it. The gain came mostly from the AI widening the list of possibilities Does LLM assistance help clinicians build better differentials?. Above that, AMIE took histories from 100 real urgent care patients without a single safety stop Can conversational AI safely take patient histories without supervision?. That is real patients, but the study had no comparison group, and its management plans fell short of doctors' on practicality and cost. The strongest evidence in the collection comes from live primary care in Nairobi. Across 39,849 visits, clinicians with an LLM safety net made 16% fewer diagnostic errors and 13% fewer treatment errors than clinicians without one Can AI safety nets reduce errors in live clinical practice?. That is the shape of evidence that counts: real visits, a comparison group, and outcomes that matter.
What you might not expect is that the Nairobi authors credit the workflow, not the model. The tool ran in the background and fit how clinicians already worked, and it was actively rolled out. The same pattern shows up elsewhere. Wrapping o3 in a diagnostic orchestration layer slightly raised its accuracy and cut cost per case by about 70%, and the gain carried over to other model families Can orchestration strategies boost diagnostic AI without better models?. In therapy, a robot and a worksheet reduced distress while a chatbot running the same language model did not Why do robots outperform chatbots in therapy despite identical language models?. So the evidence you need is about the deployed system: model plus interface plus workflow plus setting. Strong model results don't automatically carry over to the clinic.
The second lesson is that what you compare against decides what the study can prove. Therapy chatbot trials that use a waitlist as the control group mostly measure the effect of having someone to talk to. ELIZA, a 1960s chatbot, matched Woebot in one such comparison Do chatbot trials against waitlists measure real therapeutic value?. The same note argues that real proof needs head-to-head trials against existing treatments, plus evidence of how the tool actually helps. The broader therapy evidence agrees: presence and contact seem to do more work than clinical technique What makes therapeutic chatbots actually work in clinical practice?.
A last trap is mistaking clinician approval for patient benefit. Clinicians couldn't tell GPT-4's advice from experts' (45% accuracy, about chance) and rated it as more emotionally empathetic Can clinicians tell GPT-4 advice apart from expert advice?. Radiologists, meanwhile, rated advice lower when it was labeled as coming from AI, yet their accuracy depended only on whether the advice was correct Does labeling advice as AI change how clinicians use it?. Ratings and real effects can come apart, so surveys of what clinicians think can't stand in for outcomes. The collection has no randomized trial that tracks hard patient outcomes like recovery or mortality over time. That gap is the honest answer to what proof is still missing.
Sources 11 notes
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
An LLM-based diagnostic system called AMIE exceeded primary care physician performance in text-based simulated consultations across 149 case scenarios, scoring higher on 28 of 32 specialist-rated dimensions. The advantage lay in inference from gathered information rather than in eliciting history.
In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.
A single-arm study found that AMIE, a conversational AI system, conducted real clinical histories from 100 patients without requiring a single safety intervention by human supervisors. Patient attitudes toward AI improved after the interaction, though management plans trailed physicians on practicality and cost.
In 39,849 clinic visits, clinicians with access to an LLM safety net made 16% fewer diagnostic errors and 13% fewer treatment errors than those without. The authors attribute these gains to asynchronous, interface-optimized design and active deployment strategies, not model capability alone.
Show all 11 sources
On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.
A 15-day study with 38 students found that robots and worksheets significantly reduced psychological distress while a chatbot using the same LLM did not. The active ingredient was the medium—social presence and structured format—not language capability.
Comparing therapeutic chatbots to waitlist or psychoeducation controls creates false efficacy claims by measuring conversational contact rather than therapy-specific mechanisms. ELIZA matching Woebot performance demonstrates this; real evidence requires comparative trials against existing treatments and mechanism identification.
Evidence shows embodied agents and basic conversation outperform chatbots using identical clinical techniques, while LLMs struggle with core therapeutic skills like reflective listening. Physical presence and expressive contact appear to be the primary active ingredients over CBT-specific content.
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards Conversational Diagnostic AI
- Sequential Diagnosis with Language Models
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Clinical knowledge in LLMs does not translate to human interactions
- A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
- Towards Accurate Differential Diagnosis with Large Language Models
- Superhuman performance of a large language model on the reasoning tasks of a physician
- Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers