Expert advice doesn't reliably make doctors more accurate: it amplifies whatever that advice gets right or wrong, for better or worse.
How much diagnostic accuracy is gained when physicians receive expert advice?
This explores how much better physicians diagnose when someone (a human expert or an AI presented as one) gives them advice, and what decides whether that advice helps or hurts.
This explores how much physicians' diagnostic accuracy improves when they get outside advice, whether from a human expert or an AI. The corpus has no single number for it. Advice does not reliably add accuracy. It amplifies whatever the advice already contains: when the advice is right, physicians improve, and when it is wrong, they get worse, often by more than they gained.
The clearest gain comes from LLM assistance on hard cases. Twenty clinicians working through 302 difficult NEJM cases raised their top-10 differential accuracy from 36.1% to 51.7% with an LLM, compared with search alone. The authors credit the LLM's wider range of possibilities rather than sharper reasoning Does LLM assistance help clinicians build better differentials?. So the advice helped mostly by bringing candidate diagnoses into view that the physicians hadn't thought of. Against that, an experiment with professional radiologists found that AI predictions did not improve their performance on average. The radiologists gave the AI's output too little weight, and they wrongly treated their own judgment as independent of it, so the gains that combining human and AI should produce never showed up Why don't radiologists benefit from AI predictions?.
The losses can be large. In mammography, wrong BI-RADS category suggestions (BI-RADS is the standard scale for rating breast-imaging findings) cut experienced radiologists' accuracy from 82% to 45.5%. Inexperienced readers fell from nearly 80% to below 20% How much does wrong AI advice harm radiologist accuracy?. Pathology experts were misled by wrong advice in only about 7% of assessments, but time pressure made those errors much worse Does time pressure make AI advice more persuasive to experts?. The less experienced the reader and the more rushed the setting, the more the advice decides the outcome.
The label on the advice turns out to matter less than you might expect. When radiologists were told the same advice came from AI rather than a human expert, they rated it lower, but their accuracy still followed whether the advice was correct Does labeling advice as AI change how clinicians use it?. In a separate study, clinicians preferred advice they believed came from an expert 93.55% of the time, even though their guesses about who wrote it were at chance level Does the label on advice shape how clinicians judge it?. In blinded ratings, they couldn't tell GPT-4's advice from expert advice and rated GPT-4 slightly higher on empathy Can clinicians tell GPT-4 advice apart from expert advice?. So "expert" changes how much physicians trust the advice, not how much it helps them.
This leads to an awkward question. If the advisor is often stronger than the physician, a team made of a human plus advice may be the weaker setup. An LLM outperformed hundreds of physicians on differential diagnosis and triage Can language models reason better than physicians at diagnosis?. A diagnostic AI called AMIE beat primary care physicians on 28 of 32 specialist-rated measures, and its edge came from drawing conclusions from the information it had gathered Can an AI system diagnose better than primary care doctors?. A structured workflow built around o3 reached about 80% accuracy on NEJM cases at roughly a third of the cost of o3 alone Can orchestration strategies boost diagnostic AI without better models?. So the payoff from advice depends less on how good the advisor is and more on whether physicians can tell when to defer and when to override. None of these studies has solved that calibration problem.
Sources 10 notes
In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.
An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.
A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.
Among 28 pathology experts, AI-induced errors occurred in 7% of assessments regardless of time pressure, but time constraints made those errors more severe—experts relied more heavily on wrong AI advice and showed sharper performance declines.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
Show all 10 sources
Clinicians preferred advice they believed was expert-written 93.55% of the time, even though their guesses about authorship were at chance level. Their scores for quality and empathy shifted based on perceived author, not the text's actual origin.
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
An LLM-based diagnostic system called AMIE exceeded primary care physician performance in text-based simulated consultations across 149 case scenarios, scoring higher on 28 of 32 specialist-rated dimensions. The advantage lay in inference from gathered information rather than in eliciting history.
On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Towards Conversational Diagnostic AI
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Sequential Diagnosis with Language Models
- Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology
- Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
- Clinical knowledge in LLMs does not translate to human interactions