INQUIRING LINE

Giving doctors an AI's opinion often doesn't make their diagnoses better — even when the AI itself is right. Why not?

Why do radiologists fail to benefit from AI decision support?

This explores why giving radiologists (and clinicians like them) an AI's prediction often fails to make them more accurate, even when the AI itself is good, and what the collection suggests about the human side of that partnership.


This explores why giving radiologists an AI's prediction often fails to make them more accurate, even when the AI is good on its own. The short answer from the corpus is that the problem is rarely the model. It is how people fold a machine's opinion into their own judgment. In a study of professional radiologists, AI predictions did not improve average performance. The reason wasn't bad AI. Radiologists gave the AI's output too little weight, and they treated their own reading and the AI's signal as if the two were independent. That belief-updating mistake meant the potential gains from combining human and machine never showed up Why don't radiologists benefit from AI predictions?.

The puzzle is that radiologists also lean on AI too much when it's wrong. When incorrect BI-RADS categories (the standard scale for rating how suspicious a mammogram finding is) were shown as AI output, experienced readers dropped from 82% to 45.5% accuracy. Inexperienced readers fell from nearly 80% to below 20% How much does wrong AI advice harm radiologist accuracy?. So clinicians can discount correct AI advice and still be dragged down by wrong advice. An ICU study with nurses shows why that combination is so damaging. Misleading predictions hurt performance by 96–120%, while correct ones helped by only 53–67% Do wrong AI predictions hurt more than right ones help?. Averaging those two numbers hides the imbalance, which is one reason 'AI helps on average' can turn into 'AI doesn't help' in practice.

There is also a slower cost that most evaluations never measure. After AI polyp detection arrived in four Polish endoscopy centers, doctors' unassisted adenoma detection fell from 28.4% to 22.4% Does AI polyp detection weaken endoscopists' unassisted performance?. Even correct AI suggestions can cost something in the moment, because they break the clinician's focus and force them to rebuild it Does AI assistance always help reasoning or does it carry hidden costs?. A support tool can look fine in a single trial and still wear down the skills it was meant to support.

The more hopeful part of the corpus suggests the failure has a lot to do with the format of the help: a verdict handed over for the human to accept or reject. 'Learning to Guide' replaces the AI's answer with interpretive guidance, meaning it points to the parts of the input worth looking at. That is designed to remove the anchoring bias that a ready-made answer creates Can AI guidance reduce anchoring bias better than AI decisions?. Similarly, assistants that pair advice with reflection questions beat assistants that only give answers Do reflection questions help people make better decisions with AI?. And HealthBench found that physicians working with a frontier model matched or beat that model alone, so the benefit depends on who uses the AI and how Do AI models outperform physicians on health tasks?.

The takeaway you may not have expected: 'human plus AI' isn't automatically better than either one alone. Getting a gain requires the human to judge correctly when the machine knows something they don't. People are poor at that calibration, and the usual accuracy metrics hide the failure. The fix may lie less in better models than in AI that sharpens what the clinician notices rather than telling them what to conclude.


Sources 8 notes

Why don't radiologists benefit from AI predictions?

An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.

How much does wrong AI advice harm radiologist accuracy?

A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.

Do wrong AI predictions hurt more than right ones help?

In an ICU simulation, misleading AI predictions degraded nurse performance by 96–120%, while correct predictions improved it by only 53–67%. This asymmetry was hidden by standard metrics that average gains and losses together.

Does AI polyp detection weaken endoscopists' unassisted performance?

A Polish study of four endoscopy centers found adenoma detection rates fell from 28.4% to 22.4% in standard colonoscopies performed after clinicians began using AI assistance, suggesting continuous AI exposure may impair unaided diagnostic performance.

Does AI assistance always help reasoning or does it carry hidden costs?

Well-intentioned AI suggestions can damage reasoning performance by severing cognitive immersion, forcing users to rebuild focus before continuing. Evaluation must measure flow preservation across entire tasks, not just local suggestion accuracy.

Show all 8 sources
Can AI guidance reduce anchoring bias better than AI decisions?

Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.

Do reflection questions help people make better decisions with AI?

A lab study of 80 participants found that thinking assistants combining reflection questions with advice significantly outperformed agents that only advised, only questioned, or did neither. Prioritizing Socratic questioning over authoritative answers enhanced cognitive outcomes.

Do AI models outperform physicians on health tasks?

HealthBench's evaluation of 5,000 multi-turn health conversations found frontier models scored higher than physicians working alone, but physicians matched or exceeded model performance when assisted by that same model, suggesting AI benefits depend on who deploys it.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.