Do AI models outperform physicians on health tasks?
HealthBench tested whether large language models produce better health responses than physicians using a shared rubric. The question matters because it determines whether AI can replace or meaningfully augment clinical decision-making.
HealthBench, an open-source benchmark built from 5,000 multi-turn health conversations, used physician-written rubrics to grade both AI models and physicians on the same conversations. Developed with 262 physicians who have practiced in 60 countries and across 26 specialties, the benchmark scores responses against 48,562 unique rubric criteria worth between −10 and 10 points each, graded by a model-based grader the authors validated against physician judgment. Measured across roughly two years of model releases, HealthBench scores show "steady initial progress (compare GPT-3.5 Turbo's 16% to GPT-4o's 32%) and more rapid recent improvements (o3 scores 60%)." The authors also report that "GPT-4.1 nano outperforms GPT-4o and is 25 times cheaper." Physicians were asked to produce responses both with and without model assistance, and the paper's human baseline finds "recent models produce higher quality responses than physicians unless physicians are assisted by the same models."
The comparison works because physicians' responses were graded on exactly the same rubric criteria as model responses: criteria are written per-conversation by a physician, cover facts to include, communication behaviors, and context-seeking, and are summed and normalized into a 0–1 score. A subset of 34 "consensus criteria," assigned to a conversation only when two or more reviewing physicians agree it applies, lets the authors isolate narrow behaviors such as prompt escalation of emergencies. That shared rubric is what makes the physician-versus-model and physician-plus-model comparisons commensurable: unassisted physicians and the frontier models answered under identical scoring rules, and the paper reports physicians closing the gap specifically when assisted by the same model they are being compared against.
This sits in tension with Why do LLMs fail when users interact with them?, where lay users paired with an LLM did no better than unaided controls on diagnostic identification — HealthBench's physicians, by contrast, match or exceed the model's own score once assisted by it, suggesting the benefit of AI assistance on health tasks may hinge on who is doing the assisting, not only on model capability. It also extends Can language models reason better than physicians at diagnosis?: both report models outscoring physicians working alone, but HealthBench specifies the condition under which that gap closes rather than only flagging superhuman performance as awaiting trial. And it runs alongside Can conversational AI safely take patient histories without supervision?, a reminder that rubric-graded response quality and workflow-level practicality are different measures that can diverge.
The excerpt does not give the actual assisted-physician scores or a statistical comparison against model-alone scores, so the size of the "unless assisted" effect is unclear from this passage alone. The authors caution that physician agreement on consensus-criteria grading ranges only from 55% to 75%, that criteria written by individual physicians were not independently validated, and that HealthBench "does not measure health outcomes of specific workflows, which depend not only on the quality of model responses but also, critically, implementation." The implication is that rising benchmark scores and a closing physician-assistance gap are evidence of improving response quality under a rubric — not yet evidence that deploying a model at this score level changes real-world health outcomes.
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do clinicians calibrate trust in AI medical recommendations?- How much do physician scores improve when assisted by the same model?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Can rubric-graded response quality predict real-world clinical workflow success?
- Do consensus criteria identify behaviors where physicians and models differ most?
- Does AI change clinician cognition or just increase reliance on predictions?
- Why do radiologists fail to benefit from AI decision support?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do LLMs fail when users interact with them?
Standard benchmarks show LLMs excel at medical diagnosis alone, yet real users get no benefit. This explores where the breakdown happens between model capability and human decision-making.
contrasts: lay users with AI assistance didn't improve, while HealthBench's physicians matched the model once assisted by it
-
Can language models reason better than physicians at diagnosis?
An LLM was tested on challenging clinical cases against hundreds of physicians as a baseline. The research asks whether AI can match or exceed human diagnostic reasoning in structured medical settings.
both find models beating unaided physicians; HealthBench specifies when physician assistance closes that gap
-
Can conversational AI safely take patient histories without supervision?
A feasibility study tested whether an LLM-based system could conduct real clinical interviews with urgent-care patients without requiring safety interventions. Understanding AI safety in unsupervised clinical settings matters for potential deployment.
shows quality scores and workflow practicality are separate measures that can diverge
-
Can AI safety nets reduce errors in live clinical practice?
A study of 39,849 visits at Nairobi primary care clinics tested whether LLM decision support tools could help clinicians make fewer diagnostic and treatment errors in routine care, and what conditions made the tool effective.
Evidence for A: AI-assisted clinicians in Nairobi clinics made fewer errors, confirming physician-AI assistance helps in practice
-
Does LLM assistance help clinicians build better differentials?
A randomized study tested whether giving clinicians access to an LLM improved their diagnostic reasoning on challenging cases. Understanding this matters for evaluating AI's role in clinical decision support beyond standalone performance.
Evidence for A: LLM-assisted clinicians beat search alone on NEJM differential diagnosis, supporting AI assistance raising physician performance
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Clinical knowledge in LLMs does not translate to human interactions
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Sequential Diagnosis with Language Models
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Superhuman performance of a large language model on the reasoning tasks of a physician
- Towards Conversational Diagnostic AI
- A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic
Original note title
HealthBench finds models produce higher-quality health responses than physicians unless physicians are assisted by the same model