Give an LLM a complete written case and it can beat many doctors at diagnosis; trouble starts when the job reaches beyond the page.
How do LLM performances compare across different types of medical tasks?
This explores whether LLMs are uniformly good or bad at medicine, or whether their performance depends on the kind of medical task (diagnosing a written case, helping a doctor, talking to a patient, labeling records, judging clinical text), and what explains the differences.
This explores whether LLM medical performance is one number or many, and what makes some medical tasks go well while others fail. The corpus points to a clear pattern. Models do best when all the information is already on the page and the job is to reason from it. They do much worse when the task depends on what happens around the model: who supplies the information, who acts on the answer, and who has to define what a correct answer is.
At the strong end, diagnosis from a written case looks remarkable. In vignette experiments judged by physicians, and in an emergency-room study, one LLM beat hundreds of physicians on differential diagnosis, triage and management reasoning Can language models reason better than physicians at diagnosis?. AMIE, a diagnostic system tested in simulated text consultations, scored higher than primary care doctors on 28 of 32 axes rated by specialists Can an AI system diagnose better than primary care doctors?. The detail worth noticing is where its advantage came from: drawing conclusions from information already gathered, not asking the patient the right questions. The same strength helps when a doctor stays in charge. Clinicians working through 302 hard NEJM cases built better differential lists with an LLM than with search alone (51.7% vs 36.1% top-10 accuracy), mostly because the model widened the range of possibilities they considered Does LLM assistance help clinicians build better differentials?.
The most surprising result is what happens when ordinary people use these same models. In a UK trial with 1,298 participants, the models alone identified the right condition 94.9% of the time. People using them got it right only 34.5% of the time, no better than a control group without the models Why do LLMs fail when users interact with them?. The model's knowledge was fine. What broke down was the handoff: what people told the model, and how they read and acted on its suggestions. Put this next to the AMIE finding and a theme appears. The parts of medicine that run on conversation, such as drawing out a history or getting a layperson to act on a suggestion, are where performance falls apart. Benchmarks tend to miss this. Work outside medicine shows that scores on short, single-turn tests don't predict how a model holds up over long, multi-step delegated work Do short benchmarks predict how models perform over long workflows?. Single-turn tests and agent tests are really measuring two different modes of the same model Are LLM and agent benchmarks really measuring different things?.
Narrow specialist tasks show a different weakness. On clinical inference tasks, where the model judges what a piece of medical text implies, models were often wrong while sounding highly confident. Prompting tricks that improve general performance did not fix that overconfidence Why do language models fail confidently in specialized domains?. Annotation exposed yet another problem. Cancer researchers knew exactly how to label complex concepts in clinical notes by hand, yet ran into barriers getting LLMs to do it in 12 of 14 tasks. The trouble wasn't mainly accuracy. It was hard to put the concept into words precisely enough, to judge whether the notes themselves were reliable, and to check the output efficiently Can clinical experts teach LLMs to annotate complex medical concepts?. When the expertise is tacit, the bottleneck shifts from what the model knows to what the human can explain.
There is also a less obvious category: tasks where pattern-blending is the point. LLMs can simulate patients with specific thought patterns realistically enough to train therapists, rated more realistic than plain GPT-4 Can structured cognitive models improve LLM patient simulations for therapy training?. In neuroscience, fine-tuned models predicted which experimental results actually occurred better than human experts. The same tendency that produces hallucination when looking up facts becomes useful when predicting results Can LLMs predict novel scientific results better than experts?. So the useful question isn't "how good are LLMs at medicine?" It's whether a given task rewards inference from complete information or plausible extrapolation (strengths), or depends on eliciting information, calibrated confidence, and a human acting on the output (current weak points).
Sources 10 notes
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
An LLM-based diagnostic system called AMIE exceeded primary care physician performance in text-based simulated consultations across 149 case scenarios, scoring higher on 28 of 32 specialist-rated dimensions. The advantage lay in inference from gathered information rather than in eliciting history.
In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.
A trial of 1,298 UK participants found GPT-4o, Llama 3, and Command R+ scored 94.9% accuracy alone but users achieved only 34.5% when identifying conditions, no better than controls. The gap lies in how users interpret and act on model suggestions, not model knowledge.
DELEGATE-52 evaluated models across 50-round-trip relays and found short-interaction performance does not predict sustained delegation accuracy. Models ranking similarly on single-turn tasks diverged dramatically by relay 25, revealing degradation curves invisible to standard benchmarks.
Show all 10 sources
DELEGATE-52 shows that LLM benchmarks and agent benchmarks test the same model under different operating conditions—completion mode versus tool-using mode—not separate artifacts. Honest evaluation requires assessing both modes to characterize actual capability.
LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.
Cancer researchers knew exactly where and how to annotate complex concepts by hand, yet encountered barriers replicating that expertise with LLMs in twelve of fourteen tasks. The friction arose not from model accuracy alone, but from difficulties specifying concepts, judging note reliability, and evaluating results efficiently.
PATIENT-Ψ integrates 106 Beck CCD-based cognitive models with LLMs to simulate patients with specific maladaptive patterns. Expert evaluators rated the fidelity higher than GPT-4, particularly for maladaptive cognitions and conversational authenticity.
BrainBench benchmarks show fine-tuned LLMs outperform neuroscience experts at predicting which experimental results actually occurred. The same pattern-integration tendency that causes hallucination in retrieval tasks enables genuine prediction in forward-looking scenarios.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards Accurate Differential Diagnosis with Large Language Models
- Capabilities of Gemini Models in Medicine
- Clinical knowledge in LLMs does not translate to human interactions
- Sequential Diagnosis with Language Models
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
- Superhuman performance of a large language model on the reasoning tasks of a physician
- Diagnostic Reasoning Prompts Reveal the Potential for Large Language Model Interpretability in Medicine
- AI-based Clinical Decision Support for Primary Care: A Real-World Study