Does LLM assistance help clinicians build better differentials?
A randomized study tested whether giving clinicians access to an LLM improved their diagnostic reasoning on challenging cases. Understanding this matters for evaluating AI's role in clinical decision support beyond standalone performance.
The paper's central claim is that an LLM tuned for diagnostic reasoning helps clinicians build a differential diagnosis (DDx), not only that it can produce one. Twenty clinicians worked through 302 challenging NEJM case reports, each read by two clinicians randomized to either search and standard medical resources, or those tools plus the LLM. Standalone, the LLM's top-10 accuracy was 59.1% against 33.6% for unassisted clinicians (p = 0.04). In the assisted arms, the excerpt reports top-10 accuracy of 51.7% for clinicians with the LLM, against 36.1% for clinicians without its assistance (McNemar's test 45.7, p < 0.01) and 44.4% for clinicians with search (4.75, p = 0.03). The authors also say LLM-assisted clinicians reached more comprehensive differential lists. These are the authors' own measurements.
The model is PaLM 2 (large), fine-tuned with long context on medical question answering, medical dialogue and EHR note summarization. Its training data included MultiMedQA, a proprietary set of medical conversations, and expert-written MIMIC-III summaries. The authors link the long context to "tasks that require long-range reasoning and comprehension." Their proposed mechanism for the assistive effect is breadth: "the LLM's primary assistive potential may be due to making the scope of DDx more complete." The interface was pre-populated with the history of present illness, and clinicians were warned not to ask about information absent from the case, because a pilot had shown questions about lab values or imaging "leading to confabulations." That design choice keeps the dialogue inside the case as written.
Against the nearby notes, this study shifts the test from the model's answers to the clinician's output. The therapy comparison Can language models match therapist empathy in real conversations? also finds LLM strengths in single-turn responses, but this DDx study tests a different task, case workup, and measures assistance to a clinician. Where Can clinical experts teach LLMs to annotate complex medical concepts? reports barriers when experts try to reproduce their own work with LLMs, this reader study finds assisted clinicians outperforming the search arm. That is a contrast between two tasks rather than a contradiction. The medical fine-tuning also fits the domain-investment argument in Does medical AI need knowledge or reasoning more?, but the excerpt does not separate factual knowledge from reasoning, so it cannot tell which the gains depend on.
The excerpt does not establish clinical benefit. The authors chose challenging "zebras" rather than common conditions, and they write that their evaluation "does not directly indicate" results for typical daily cases. The model saw only the main text, while clinicians also had images and tables, and the authors cannot say how much that gap would matter. The standalone edge over GPT-4 rests on automated model-based evaluation, not clinician ratings. The authors also note that performance at clinicopathological conferences "in no way reflects a broader measure of competence," and the interviewed clinicians judged learning and education the most appropriate use at present, since "additional work is needed to understand suitability for clinical settings." The supportable reading is narrower than a clinical claim: in a randomized reader study on hard cases, LLM assistance widened and improved differential lists. Whether that improves patient diagnosis is not shown.
Inquiring lines that read this note 16
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do clinicians calibrate trust in AI medical recommendations?- How much does self-play training with LLM-simulated patients actually improve diagnostic accuracy?
- Why does medical knowledge require continuous access to current sources?
- What evidence would prove medical AI actually works in clinics?
- How does blinded rating of diagnoses compare to real clinical outcomes?
- What role does cost estimation play in steering diagnostic test ordering?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Does this colonoscopy finding apply to other medical specialties using AI?
- How much do physician scores improve when assisted by the same model?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Does AI change clinician cognition or just increase reliance on predictions?
- Can offline LLM evaluation predict performance in live clinical workflows?
- How much does missing images and tables limit LLM diagnostic reasoning?
- Can LLM performance on zebra cases predict results in routine clinical practice?
- Does medical fine-tuning help LLMs through knowledge or reasoning ability?
- How do LLM performances compare across different types of medical tasks?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can language models match therapist empathy in real conversations?
Do LLMs' high empathy scores on isolated responses translate to therapeutic skill in actual ongoing treatment? This explores whether single-turn advantage predicts real-world therapeutic performance.
parallel clinical evaluation of single-turn LLM strengths; this study tests case workup and clinician assistance, not therapy
-
Can clinical experts teach LLMs to annotate complex medical concepts?
Clinical experts can manually identify complex medical concepts in patient notes, but transferring that expertise to LLM-based extraction systems proves difficult. Understanding where this transfer breaks down could improve how AI tools support expert workflows.
contrast: experts there met barriers reproducing their work with LLMs, while here assisted clinicians beat the search arm; the tasks differ
-
Does medical AI need knowledge or reasoning more?
Medical and mathematical domains may require fundamentally different AI training priorities. If medical accuracy depends primarily on factual knowledge while math depends on reasoning quality, should we build and evaluate these systems differently?
qualifies: medical fine-tuning fits a domain-investment argument, but the excerpt does not separate knowledge from reasoning, so the asymmetry stays untested
-
Why do LLMs fail when users interact with them?
Standard benchmarks show LLMs excel at medical diagnosis alone, yet real users get no benefit. This explores where the breakdown happens between model capability and human decision-making.
qualifies: a 1,298-person UK trial found LLM assistance gave public users no gain over controls, against the clinician-reader gain
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Towards Accurate Differential Diagnosis with Large Language Models
- Capabilities of Gemini Models in Medicine
- Clinical knowledge in LLMs does not translate to human interactions
- Diagnostic Reasoning Prompts Reveal the Potential for Large Language Model Interpretability in Medicine
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
- Sequential Diagnosis with Language Models
- Superhuman performance of a large language model on the reasoning tasks of a physician
Original note title
LLM assistance lifted top-10 differential accuracy above search on NEJM cases — the authors credit wider lists