When AI offers a scan reading, how do radiologists keep their own judgment from giving way to it?
What safeguards help radiologists maintain independent judgment when using AI assistance?
This explores what keeps a radiologist's own reading from collapsing into the AI's suggestion, and what the collection says about how independent judgment gets lost in the first place.
This explores what keeps a radiologist's own reading from collapsing into the AI's suggestion. The collection doesn't have a tested checklist of radiology safeguards. What it does have is a clear account of how independent judgment fails, plus a few design ideas from other fields that point toward fixes. The failure is worse than most people expect. In a mammography study, wrong AI suggestions dropped experienced radiologists from 82% to 45.5% accuracy, and inexperienced readers fell below 20% How much does wrong AI advice harm radiologist accuracy?. Experience softened the damage but didn't prevent it.
The surprising part is that 'stay independent' isn't simply the right goal. A separate experiment found that AI predictions don't improve radiologists on average, partly because they give the AI too little weight. They also treat their own read and the AI's output as two independent opinions when the two often draw on the same visual evidence Why don't radiologists benefit from AI predictions?. So radiologists defer too much when the AI is wrong and combine signals badly when it's right. A useful safeguard has to fix calibration, not just resistance. One promising sign: in that same study, giving radiologists contextual clinical information helped them where the AI prediction alone did not.
The strongest design idea comes from outside radiology. 'Learning to Guide' proposes that the AI shouldn't hand over a verdict at all. Instead it highlights which parts of the input deserve attention and leaves the decision with the human. The authors report that this removes anchoring, the pull of a suggested answer Can AI guidance reduce anchoring bias better than AI decisions?. A related idea comes from work on guarding automated judges: run the checks that can't be argued with before the ones that can, and plant known test cases as alarms Can deterministic checks protect LLM judges from failure?. Applied to a reading room, that suggests two practices: record your own read before the AI's suggestion appears, and slip known-answer cases into the workflow to show whether readers are rubber-stamping. That's an inference from the collection, not a tested radiology protocol.
There's also a slower threat that per-case safeguards miss. After AI polyp detection arrived in Polish endoscopy centers, adenoma detection in colonoscopies done without AI fell from 28.4% to 22.4% Does AI polyp detection weaken endoscopists' unassisted performance?. Independent judgment is a skill that weakens with disuse, so one safeguard is regular practice without AI, with the results measured. The broader pattern has a name, 'cognitive surrender': people stop checking because checking costs effort and fluent output feels trustworthy When do users stop checking whether AI output is actually backed?. The systems most likely to cause this are the ones that seem most competent How do competent systems quietly undermine safety oversight?.
Two things work against the obvious fixes. First, interrupting someone to make them reconsider has a cost. AI interventions can break concentration even when they're correct Does AI assistance always help reasoning or does it carry hidden costs?, so the timing of a prompt matters as much as what it says. Second, people using AI expect colleagues to see them as less competent and tend to hide that they used it Do people fear judgment when they use AI at work?. Any safeguard that relies on clinicians openly reporting when they leaned on the AI will run into that.
Sources 9 notes
A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.
An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.
Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.
Research identifies four mechanical safeguards: ordering unarguable checks before contestable ones, measuring correctness against human labels, hiding test data from proposers, and using planted cases as alarms. None requires the LLM itself to verify compliance.
A Polish study of four endoscopy centers found adenoma detection rates fell from 28.4% to 22.4% in standard colonoscopies performed after clinicians began using AI assistance, suggesting continuous AI exposure may impair unaided diagnostic performance.
Show all 9 sources
Users systematically accept AI outputs without verification because checking is costly and fluent output builds false confidence. This receiver-side surrender—measured in studies showing 80% unchallenged adoption—is what enables inflationary token systems to function at scale.
The most dangerous AI systems appear to function well while weakening skepticism through fluent outputs, collapsing authority boundaries by treating context as instruction, storing unsafe state across time in workflows, and diffusing accountability across multiple actors. Evidence includes overconfident model outputs, prompt injection payloads bypassing guards, and poisoned shared memory in multi-agent pipelines.
Well-intentioned AI suggestions can damage reasoning performance by severing cognitive immersion, forcing users to rebuild focus before continuing. Evaluation must measure flow preservation across entire tasks, not just local suggestion accuracy.
Across four experiments with 4,439 participants, people using AI expected others to judge them as less competent and diligent, and reported lower willingness to disclose AI use to managers and colleagues. The gap suggests a social cost that users foresee and act on.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Automation Bias in AI-Assisted Medical Decision-Making under Time Pressure in Computational Pathology
- How AI Can Degrade Human Performance in High-Stakes Settings
- Humans learn to prefer trustworthy AI over human partners
- Epistemic Deference to AI