INQUIRING LINE

Do radiologists working with AI have an accurate sense of whether it really improves their diagnoses, or do they misjudge it?

Do radiologists' beliefs about AI-assisted performance match their actual outcomes?

This explores whether radiologists working with AI understand how the AI is actually affecting their accuracy, or whether what they believe about the collaboration differs from what the results show.


This explores whether radiologists working with AI have an accurate sense of how the AI changes their diagnoses. The short answer from the corpus is no, and the mismatch goes in two directions at once. No study here simply asks radiologists to predict their AI-assisted accuracy and then checks it. What the corpus does show is a pattern of miscalibration that points clearly toward a gap between belief and outcome.

The first direction is trusting the AI too little. In an experiment with professional radiologists, AI predictions did not raise average performance, even though the AI carried useful signal Why don't radiologists benefit from AI predictions?. The researchers trace this to errors in how radiologists update their beliefs. They gave AI output less weight than it deserved, and they treated their own reading as independent evidence when it actually overlapped with what the AI had already seen. In practice, radiologists acted as if they knew better than the combined evidence did, and that is exactly the kind of gap between belief and result the question asks about.

The second direction is trusting the AI too much, and it shows up at the moment the AI is wrong. In a 27-radiologist mammography study, incorrect AI-labeled BI-RADS categories (BI-RADS is the standard scale radiologists use to grade how suspicious a breast finding is) pulled experienced readers from 82% accuracy down to 45.5%. Inexperienced readers fell from nearly 80% to below 20% How much does wrong AI advice harm radiologist accuracy?. Put this next to the first study and you get a paradox: clinicians underuse AI on average but follow it when it misleads them. One likely reason comes from outside medicine. Users in every language studied follow how confident the AI sounds rather than whether it is right Do users worldwide trust confident AI outputs even when wrong?. A confident wrong label carries more weight than a correct one that looks routine.

The most unsettling evidence is about skills slowly eroding without anyone noticing. After four Polish endoscopy centers adopted AI polyp detection, adenoma detection in colonoscopies done without AI fell from 28.4% to 22.4% (adenomas are precancerous growths, so missing them matters) Does AI polyp detection weaken endoscopists' unassisted performance?. Nothing in that study suggests the clinicians felt any less skilled. Two findings from general AI use explain why they might not. People tend to count AI-assisted output as proof of their own ability, especially when the collaboration feels seamless Do AI-assisted outputs fool users about their own skills?. And across three pooled studies, people's self-rated AI competence correlated with their measured competence at only .055, which is effectively zero Can self-ratings replace objective performance scores for AI competence?.

The takeaway you might not have expected: the fix probably isn't better intuition but a better track record. A model-side result offers a useful parallel. AI confidence becomes reliable when it is anchored to how similar past cases actually turned out, not to how the current case feels Can past performance predict when a model will be right?. If radiologists regularly saw their own accuracy with and without AI, case by case, the gap between what they believe and what actually happens would become something they could see and correct.


Sources 7 notes

Why don't radiologists benefit from AI predictions?

An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.

How much does wrong AI advice harm radiologist accuracy?

A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.

Do users worldwide trust confident AI outputs even when wrong?

Cross-linguistic research shows users in every language trust confident AI outputs even when inaccurate. While confidence expression varies by language, users everywhere track confidence signals rather than accuracy, making overconfident errors systematically followed.

Does AI polyp detection weaken endoscopists' unassisted performance?

A Polish study of four endoscopy centers found adenoma detection rates fell from 28.4% to 22.4% in standard colonoscopies performed after clinicians began using AI assistance, suggesting continuous AI exposure may impair unaided diagnostic performance.

Do AI-assisted outputs fool users about their own skills?

Research identifies a systematic cognitive attribution error where individuals integrate AI-generated outputs into their capability identity, believing they possess skills they don't actually have. This occurs when task output is seamless and fluent, obscuring the human-AI boundary.

Show all 7 sources
Can self-ratings replace objective performance scores for AI competence?

A pooled analysis of three studies found a correlation of only .055 between self-reported and objective measures of AI competence, with confidence intervals including zero. This provides no basis for substituting self-assessment for demonstrated performance.

Can past performance predict when a model will be right?

XConf matches ten-sample self-consistency at a tenth of the cost by retrieving the model's past episodes with similar confidence levels and reading their historical success rates. Ablations show the signal depends entirely on stored outcomes, not on the retrieval prompt itself.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.