Why don't radiologists benefit from AI predictions?
When radiologists receive AI predictions, they often fail to incorporate them properly into their decisions. This explores what belief-updating errors prevent radiologists from realizing potential AI-assisted gains.
Agarwal, Moehring, Rajpurkar and Salz run an information experiment with professional radiologists and report a split result. Providing AI predictions "does not improve performance on average," whereas providing contextual information, meaning information the radiologists hold that the AI does not, does. The sharper claim concerns the shortfall. Radiologists "do not realize the gains from AI assistance," and the authors attribute this to "errors in belief updating." The excerpt treats the gains as real but unrealized.
The mechanism is an updating account, not a missing-information account. The excerpt names two errors: radiologists "underweight AI predictions," and they "treat their own information and AI predictions as statistically independent." The second error is the one that bears on combination. Two signals that share information cannot be weighed as if they were independent sources, so a clinician who assumes independence will merge them wrongly. The loss, on this account, sits at the join between the human's knowledge and the AI output, not in either source alone. The humans are not described as lacking information; the AI is described as lacking the contextual information the humans already have.
Set against the nearest notes, this excerpt locates a gap they describe at a different point. Does theory of mind predict who thrives in AI collaboration? finds that collaborative performance separates from solo performance and is predicted by perspective-taking. This excerpt points at the step after the user reads the AI output, which is how that output is weighed against the user's own view. The sharpest tension is with Can AI guidance reduce anchoring bias better than AI decisions?. That note prefers guidance to deferral. The excerpt's design conclusion cuts the other way: "the optimal human-AI collaboration design delegates cases either to humans or to AI, but rarely to AI assisted humans," unless the belief-updating mistakes can be corrected. Guidance only informs a decision if the human updates on it, and this excerpt says radiologists do not do so reliably. Can self-ratings replace objective performance scores for AI competence? is a parallel case in which what users believe about their AI-assisted performance does not track measured outcomes, though that note measures self-report, not belief updating.
The excerpt does not establish how large the gap is. It is an abstract, with no effect sizes, sample size, radiology task, case mix, or description of how belief updating was tested, so the direction of the independence error is not shown. It also does not show that AI assistance makes radiologists worse; it shows that AI predictions do not improve average performance in this experiment and attributes the shortfall to updating. The implication, at the strength the abstract supports, is narrow. Before assuming an AI output adds value in a human review workflow, one should check whether the human updates on it, and the paper's own conclusion conditions AI-assisted rollout on correcting those errors. Whether that correction is feasible is the open question the excerpt leaves.
Inquiring lines that read this note 13
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do clinicians calibrate trust in AI medical recommendations?- Do radiologists' beliefs about AI-assisted performance match their actual outcomes?
- Does showing AI confidence scores reduce radiologist over-reliance on wrong suggestions?
- What safeguards help radiologists maintain independent judgment when using AI assistance?
- Why do clinicians fail to act on correct AI suggestions in real care?
- Why do experts resist AI recommendations that contradict their own judgments?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Does this colonoscopy finding apply to other medical specialties using AI?
- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Why do nurses misclassify emergencies differently with misleading AI assistance?
- Does AI change clinician cognition or just increase reliance on predictions?
- Why do radiologists fail to benefit from AI decision support?
- How does trusting wrong AI advice change what medical action people decide to take?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does theory of mind predict who thrives in AI collaboration?
Explores whether perspective-taking ability—the capacity to model another's cognitive state—differentiates humans who benefit most from working with AI, separate from solo problem-solving skill.
same collaborative-versus-solo gap, located here in belief updating rather than perspective-taking
-
Can AI guidance reduce anchoring bias better than AI decisions?
When humans and AI collaborate on decisions, does providing interpretive guidance instead of proposed answers reduce both over-trust in machines and abandonment on hard cases?
contrast: the optimal design here delegates whole cases, which cuts against guidance unless the human updates on it
-
Can self-ratings replace objective performance scores for AI competence?
Do people's perceptions of their own AI competence match what they can actually do? This matters because assessment systems might rely on the wrong type of measure to evaluate workplace readiness.
parallel belief-versus-outcome gap, measured through self-report rather than belief updating
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
- How AI Can Degrade Human Performance in High-Stakes Settings
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Clinical knowledge in LLMs does not translate to human interactions
- Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology
- Epistemic Deference to AI
Original note title
AI predictions do not improve radiologists on average while contextual information does — belief-updating errors keep the AI gain unrealized