INQUIRING LINE

A medical AI beats doctors solo — so why don't doctors using that same AI do even better?

Why did lay users with AI models fail to match unaided physicians on diagnosis?

This explores why ordinary people using an AI model didn't diagnose as well as doctors working without AI, even though the model alone often beats doctors. The corpus has no study that directly tests lay users with AI, so this answer pieces the explanation together from nearby findings.


This explores why ordinary people using an AI model didn't diagnose as well as doctors working without AI, even though the model alone often beats doctors. One caveat first: none of the studies in this collection directly tests lay users with AI against doctors. What the corpus does have is a strong set of nearby findings, and together they point to one explanation. The model's diagnostic skill is real, but the person using it limits how much of that skill comes out.

Start with what the model can do by itself. On case write-ups and emergency-room cases, an LLM outperformed hundreds of physicians at diagnosis and triage Can language models reason better than physicians at diagnosis?. In simulated consultations, the diagnostic system AMIE beat primary care doctors on 28 of 32 measures Can an AI system diagnose better than primary care doctors?. AMIE's study adds a useful detail: its edge came from *reasoning about the information it had*, not from *asking the patient the right questions*. That points to the weak spot. In a vignette, the facts are already written out and complete. In a real conversation, a layperson has to know which symptoms matter, describe them well, and ask good follow-up questions. If they leave out the detail that matters, even an excellent reasoner can't use it.

The models also don't track what they haven't been told. Assistants have no internal sense of what they don't yet know about the person they're talking to. When researchers added an explicit list of 'still unknown' facts to the prompt, harmful advice and sycophancy (telling users what they want to hear) fell by 50–75% Do language models know what they don't know about users?. A layperson won't notice the gaps either. A doctor working alone knows what to ask next. A model talking with a layperson often doesn't, and the layperson can't make up for it. Confidence makes this worse. In specialized medical domains, models stay confident even when they're wrong, and prompting tricks don't fix it Why do language models fail confidently in specialized domains?.

Then there's the reader's side. People couldn't tell AI medical answers from doctors' answers (they guessed at chance level). They also rated *inaccurate* AI answers as trustworthy enough to act on Can people tell AI medical advice from doctors' responses?. So a layperson can't tell which of the model's suggestions to keep and which to drop. Compare clinicians: their top-10 diagnosis accuracy rose from 36% to 52% with LLM help, mainly because the model widened their list of possible diagnoses Does LLM assistance help clinicians build better differentials?. That works because a doctor can judge a longer list. For a layperson, a longer list is just more options they can't evaluate.

The interesting twist is that even experts don't use AI well by default. Radiologists gave AI predictions too little weight, so on average AI advice didn't improve them Why don't radiologists benefit from AI predictions?. HealthBench found that frontier models beat doctors working alone, while doctors using the same model matched or beat it Do AI models outperform physicians on health tasks?. The structure around a model matters too. Wrapping o3 in a system that orders tests and checks hypotheses step by step raised accuracy and cut costs by 70%, with no better model needed Can orchestration strategies boost diagnostic AI without better models?. Doctors supply that structure from their own training. Laypeople don't have it, and a plain chat window doesn't provide it. The lesson: good diagnosis from a human and a model together comes from how the conversation is run, not only from how smart the model is.


Sources 9 notes

Can language models reason better than physicians at diagnosis?

In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.

Can an AI system diagnose better than primary care doctors?

An LLM-based diagnostic system called AMIE exceeded primary care physician performance in text-based simulated consultations across 149 case scenarios, scoring higher on 28 of 32 specialist-rated dimensions. The advantage lay in inference from gathered information rather than in eliciting history.

Do language models know what they don't know about users?

Research shows assistants suffer from sycophancy and hallucination because they have no representation of what remains unknown about users. Adding a schema of labeled unknowns to prompts reduced harmful advice and sycophancy by 50–75% and cut hallucination rates by roughly half.

Why do language models fail confidently in specialized domains?

LLMs trained on general text lack sufficient exposure to domain-specific examples, leading to low accuracy paired with high confidence in clinical NLI tasks. Prompting techniques that improved general performance fail to reduce overconfidence in specialized domains.

Can people tell AI medical advice from doctors' responses?

A 300-participant study found participants could not reliably distinguish AI-generated medical responses from doctors' (50% accuracy, chance level) and rated low-accuracy AI answers as valid and trustworthy enough to act on them—comparable to or stronger than their trust in actual doctors' advice.

Show all 9 sources
Does LLM assistance help clinicians build better differentials?

In a study of 20 clinicians on 302 NEJM cases, those with LLM access achieved 51.7% top-10 accuracy versus 36.1% without it. The authors attribute the gain to the LLM's wider differential scope, making lists more comprehensive.

Why don't radiologists benefit from AI predictions?

An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.

Do AI models outperform physicians on health tasks?

HealthBench's evaluation of 5,000 multi-turn health conversations found frontier models scored higher than physicians working alone, but physicians matched or exceeded model performance when assisted by that same model, suggesting AI benefits depend on who deploys it.

Can orchestration strategies boost diagnostic AI without better models?

On 304 NEJM cases, MAI-DxO orchestration achieved 79.9% accuracy at $2,397 per case versus 78.6% at $7,850 for o3 alone. The gains transferred across model families, suggesting the benefit comes from the scaffold, not model weights.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.