Give doctors the same AI that already beats them solo — do they just tie it, or pull ahead of both?
How much do physician scores improve when assisted by the same model?
This explores what happens when doctors work alongside the same AI model that beats them on its own: do they just catch up, or do they end up ahead of both?
This explores what happens when doctors use the very AI model that outperforms them on its own: do they simply close the gap, or do they end up ahead of both? The clearest evidence comes from HealthBench, which graded 5,000 multi-turn health conversations. Frontier models scored higher than physicians working alone. When physicians were given that same model to work with, they matched or beat the model's own score Do AI models outperform physicians on health tasks?. The corpus doesn't give an exact point gain for that comparison. What it shows is the direction: the model's lead over doctors disappears once the doctors are the ones using it. That suggests the benefit depends heavily on who deploys the tool and how.
Where the corpus does have hard numbers, they come from diagnosis. In a study of 20 clinicians working through 302 hard New England Journal of Medicine (NEJM) case reports, those with LLM access put the right answer in their top 10 possible diagnoses 51.7% of the time, versus 36.1% with search alone Does LLM assistance help clinicians build better differentials?. That's roughly a 15-point jump. The mechanism matters: the model didn't make doctors better reasoners. It widened the list of possibilities they considered. The AI works as a brainstorming partner more than an oracle.
Here's the twist you may not have expected. Human-plus-AI teams often fail to capture the full benefit the AI offers. A 535-person study found that when an LLM got better on specific questions, the people using it captured only about half of that improvement. Sometimes they did worse than the AI would have done alone Why does assisted accuracy capture only half the LLM gain?. Having two strong partners creates potential for a better result, but it doesn't guarantee one. People override correct suggestions, accept wrong ones, or don't know when to trust which. So the HealthBench result, where physicians matched or exceeded the model, is an encouraging case, not a general rule.
The same NEJM-style cases show a different route to improvement that has nothing to do with humans at all. Wrapping a model in a structured workflow, with multiple roles that question each other and keep track of test costs, raised o3's diagnostic accuracy slightly and cut the cost per case by about 70% Can orchestration strategies boost diagnostic AI without better models?. If a well-designed process around the model can produce gains, part of what a skilled physician adds may be exactly that: structure, skepticism, and knowing which questions to ask next.
One caution before you take any of these numbers at face value: scores depend on how they're measured. Elsewhere in the corpus, apparent gains from specialized medical models largely vanished once each model got a fairly tuned prompt and proper error bars Does medical pretraining actually improve model performance?. More broadly, rising evaluation scores can hide flat real-world performance Can a higher evaluation score hide poor task performance?. The honest summary: assisted physicians reliably close the gap with the model, and in diagnosis they can gain about 15 points. Whether they consistently surpass it depends on the task, the workflow, and how carefully the scoring was done.
Sources 6 notes
HealthBench's evaluation of 5,000 multi-turn health conversations found frontier models scored higher than physicians working alone, but physicians matched or exceeded model performance when assisted by that same model, suggesting AI benefits depend on who deploys it.
In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.
A 535-participant study found that when LLM accuracy improved on individual items, assisted participants captured roughly half that gain—falling below what the better-performing component could have provided alone. This shows complementarity creates potential but does not guarantee synergy.
On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.
Seven medical LLMs and two medical VLMs showed significant improvements over base models in only 9.4% and 6.3% of tasks respectively when prompts were optimized per model and confidence intervals were reported, down from 70.5% and 62.5% without these controls.
Show all 6 sources
When systems optimize toward evaluation scores, measured progress can rise while actual task performance remains flat or declines, because optimization can exploit weaknesses in the measurement itself rather than solve the task. A relayed prompt case demonstrated this: judge pass rates rose from 23.1 to 80.0 percent while task-facing defect detection stayed unchanged.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Sequential Diagnosis with Language Models
- Clinical knowledge in LLMs does not translate to human interactions
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
- Towards Accurate Differential Diagnosis with Large Language Models
- Capabilities of Gemini Models in Medicine
- Evaluating Large Language Models in Theory of Mind Tasks