Many headline medical AI scores rest on doctor-written answer keys, so what happens when those doctors disagree with each other?
How does expert annotation instability affect medical AI benchmarking?
This explores what happens to medical AI scores when the experts who write the 'correct answers' disagree with each other or change their minds. The collection has no study that measures this directly, but it has nearby material on where medical ground truth comes from, how evaluators drift, and what disagreement can tell you.
This explores what happens to medical AI scores when the experts who write the 'correct answers' disagree with each other or change their minds. To be direct: the collection has no paper that measures how much physician labels vary or how that variation changes a benchmark's rankings. What it does have shows where the problem sits in medical AI evaluation, and why it matters more than headline numbers suggest.
Start with what headline results rest on. When an LLM is reported to beat hundreds of physicians at diagnosis, the scoring itself comes from physicians: the vignettes are 'physician-adjudicated' Can language models reason better than physicians at diagnosis?. The orchestration result on NEJM cases is similar. It reports 79.9% versus 78.6% accuracy, a gap of about one point, measured against published case answers Can orchestration strategies boost diagnostic AI without better models?. If expert adjudicators disagreed even a few percent of the time, a one-point difference could fall inside that noise. The cost savings in the same study are much less sensitive to label quality than the accuracy claim. So the takeaway is not that these results are wrong. It's that the size of the claimed effect should be read against how stable the answer key is, and those papers don't report that.
The evaluation literature gives a way to think about unstable judges. In general (non-medical) agentic tasks, LLM-as-a-Judge evaluators showed 31% 'judge shift', meaning their verdicts on the same work drifted. An agent that gathers evidence before judging got that down to 0.27% Can agents evaluate AI outputs more reliably than language models?. Human expert graders aren't LLMs, but the lesson carries over: a judge that rules without collecting evidence drifts. Medical benchmarks whose labels come from a quick expert read rather than follow-up outcomes, pathology, or consensus likely have the same weakness. A counterintuitive finding from qualitative coding points the other way: when LLM coders argued long and left some labels marked 'undecidable', their accuracy was higher Does disagreement between AI coders signal better accuracy?. That suggests disagreement among annotators may be information worth keeping. A medical benchmark that collapses three split radiologists into one 'gold' label throws away the signal that the case is genuinely ambiguous.
The less obvious risk is that labels and AI can contaminate each other. In mammography, radiologists shown a wrong BI-RADS category labeled as AI output fell from 82% to 45.5% accuracy, and less experienced readers fell below 20% How much does wrong AI advice harm radiologist accuracy?. If future benchmark annotators work with AI suggestions on screen, their 'ground truth' may partly echo the models being tested. That would make benchmarks look more stable and models look better than either really is.
Shaky labels also stack with blind spots benchmarks already have. Models are overconfident in clinical inference tasks even when they're wrong Why do language models fail confidently in specialized domains?. Training can raise final-answer accuracy while making the reasoning worse Does supervised fine-tuning improve reasoning or just answers?. And identical outputs can hide very different internal workings Can AI pass every test while understanding nothing?. A benchmark that scores only the final answer, against an answer key experts don't fully agree on, can be fooled in two ways at once. For a more trustworthy medical evaluation, look for three things: reported agreement between annotators, credit for flagging a case as ambiguous, and labels made without AI suggestions in view.
Sources 8 notes
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Multi-agent LLM coding systems showed higher accuracy when agents engaged in prolonged, unresolved debate. The frequency of disagreement and undecidable labels serve as reliable performance indicators, suggesting conflict deepens interpretive work rather than signaling failure.
A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.
Show all 8 sources
LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.
Supervised fine-tuning improves final-answer accuracy on benchmarks but cuts Information Gain by 38.9 percent, meaning models generate correct answers through post-hoc rationalization rather than genuine inferential steps. Standard metrics miss this degradation because they only measure final correctness.
The Fractured Entangled Representation hypothesis shows that SGD-trained networks can produce identical outputs across all inputs while maintaining radically different internal representations. Standard benchmarks cannot detect this structural difference.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Sequential Diagnosis with Language Models
- Towards Accurate Differential Diagnosis with Large Language Models
- Towards Conversational Diagnostic AI
- Capabilities of Gemini Models in Medicine
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Superhuman performance of a large language model on the reasoning tasks of a physician