How much does wrong AI advice harm radiologist accuracy?
When mammography radiologists receive incorrect AI suggestions labeled as system output, how much does their diagnostic accuracy decline? This matters for understanding automation bias in clinical workflows.
Dratsch et al. report, through RSNA, that radiologists reading mammograms with AI decision support became significantly worse at assigning the correct BI-RADS category when the suggestion was wrong. In a prospective experiment, 27 radiologists read 50 mammograms and gave their BI-RADS assessments with AI assistance, using two randomized sets: a training set of 10 with correct AI suggestions, and a test set of 40 in which 12 carried incorrect categories, purportedly suggested by AI. Inexperienced radiologists assigned the correct score "in almost 80% of cases" when the suggestion was correct and fell to less than 20% when it was wrong. Experienced radiologists, with more than 15 years of experience on average, dropped from 82% to 45.5%.
The excerpt reads the result as automation bias, defined as "the tendency of humans to favor suggestions from automated decision-making systems." It compares cases where the purported suggestion was right with cases where it was wrong, and the drop appears in both experience groups, though the experienced drop is smaller. The authors tie the concern to the workflow itself: "Given the repetitive and highly standardized nature of mammography screening, automation bias may become a concern when an AI system is integrated into the workflow." The excerpt notes that earlier studies of computer-aided detection had found performance impairments, but none had looked at AI systems and accurate readings. The safeguards it lists (showing confidence or per-output probabilities, teaching users how the system reasons, and keeping users accountable for their own decisions) are offered as possibilities, not tested.
Against the library, this is the measured side of a risk that Can AI guidance reduce anchoring bias better than AI decisions? describes in design terms. That note argues that deferral-style systems risk anchoring bias, where the human over-trusts the machine's decision. The mammography excerpt gives that concern an empirical case in a clinical reading task, though it names the mechanism automation bias rather than anchoring, and it does not test the guidance-style design the LTD-LTG note proposes as the fix. The closer link on the cognitive side is Does AI assistance actually harm the way developers learn?, where high-engagement interaction patterns preserved learning outcomes. That fits the excerpt's accountability safeguard, but the two studies measure different things (conceptual learning in a coding session versus accuracy on individual readings), so the connection is a hypothesis rather than a shared finding.
The excerpt is a press release, not the paper, so sample details beyond the headline counts, the statistical tests behind "significantly worse," and the per-reader spread are not given here. "Purportedly" carries the most weight: the wrong categories were labeled as AI output, so the study tests how readers respond to a labeled suggestion that is wrong, not how a deployed system's actual errors would affect them. Twenty-seven readers working through a fixed set of 50 cases is also a long way from a live screening workflow. The excerpt does not show why readers deferred either; the researchers plan eye-tracking to study that, so automation bias remains their interpretation. The implication is narrower than the headline. The result supports the authors' call for safeguards when AI enters the reading workflow, but the excerpt does not show that confidence displays, reasoning explanations or accountability prompts would reduce the effect.
Inquiring lines that read this note 15
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do clinicians calibrate trust in AI medical recommendations?- Do radiologists' beliefs about AI-assisted performance match their actual outcomes?
- Does optimizing for differential diagnosis accuracy risk pushing AI systems toward premature problem-solving?
- Does showing AI confidence scores reduce radiologist over-reliance on wrong suggestions?
- What safeguards help radiologists maintain independent judgment when using AI assistance?
- Why do clinicians fail to act on correct AI suggestions in real care?
- How does expert annotation instability affect medical AI benchmarking?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- What role does cost estimation play in steering diagnostic test ordering?
- Do physicians follow incorrect advice more when they trust its source?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Does this colonoscopy finding apply to other medical specialties using AI?
- Why do nurses misclassify emergencies differently with misleading AI assistance?
- Why do radiologists fail to benefit from AI decision support?
- Why does labeling advice as AI from a doctor change how people trust it?
- How does trusting wrong AI advice change what medical action people decide to take?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can AI guidance reduce anchoring bias better than AI decisions?
When humans and AI collaborate on decisions, does providing interpretive guidance instead of proposed answers reduce both over-trust in machines and abandonment on hard cases?
a clinical case of the over-trust risk it names; the guidance fix is untested here
-
Does AI assistance actually harm the way developers learn?
When developers use AI tools while learning new programming concepts, does it impair their ability to understand code, debug problems, and build lasting skills? Understanding this matters for how we deploy AI in education and training.
both tie AI assistance to human judgment; the engagement finding fits the accountability safeguard, but the measures differ
-
Does time pressure make AI advice more persuasive to experts?
When pathologists work under time constraints, does pressure to decide quickly make them more likely to trust and act on AI recommendations, even when those recommendations are wrong?
Evidence for, in pathology: erroneous AI advice overturned correct expert estimates in 7% of assisted assessments, showing automation bias beyond radiology
-
Does labeling advice as AI change how clinicians use it?
When physicians know diagnostic advice comes from an AI system rather than a human expert, do they rely on it differently? This matters because AI labels might trigger skepticism that affects clinical decisions.
Qualifies the AI attribution: accuracy fell with inaccurate advice whatever its source, label included
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- AI-based Clinical Decision Support for Primary Care: A Real-World Study
- Towards Conversational Diagnostic AI
- How AI Can Degrade Human Performance in High-Stakes Settings
Original note title
Dratsch et al. find wrong BI-RADS categories labeled as AI output cut accuracy for experienced radiologists from 82% to 45.5%