Does labeling advice as AI change how clinicians use it?
When physicians know diagnostic advice comes from an AI system rather than a human expert, do they rely on it differently? This matters because AI labels might trigger skepticism that affects clinical decisions.
Gaube et al. tested whether telling physicians that diagnostic advice came from "an AI system" changes how they use it. Radiologists (n = 138) and internal or emergency medicine physicians (n = 127) reviewed eight chest X-ray cases. Each case came with advice that was either accurate or inaccurate and labeled as coming from an AI system or an experienced radiologist. All of the advice was written by human experts, so only the label varied. The label moved ratings in one group: "only participants with higher task expertise showed algorithmic aversion by rating the quality of advice to be significantly lower when it came from the AI in comparison to the human." It did not move diagnoses: "the purported source of the advice did not affect participants' performance." What did move accuracy was the advice itself. Task experts performed 40.10% better, and non-experts 37.53% better, when they received accurate rather than inaccurate advice.
The paper's account of the diagnostic result is that advice pulls judgment toward itself. The authors say the advice "could have engaged cognitive biases, by anchoring participants to a particular diagnosis, and triggering confirmatory hypothesis testing," and they describe "a general tendency for participants to agree with advice," stronger among physicians with less task expertise. They define clinical susceptibility as "the propensity to follow incorrect advice." On that measure, 41.73% of IM/EM physicians and 27.54% of radiologists always gave the wrong diagnosis when the advice was wrong. Susceptibility was not limited to weaker performers: 28.26% of radiologists and 17.32% of IM/EM physicians refuted all the incorrect advice they saw. The by-source breakdown of susceptibility sits in a supplementary figure that the excerpt does not include.
This sits against the claim that labels shape reliance. Does the label on advice shape how clinicians judge it? reports a label moving clinicians' preferences. This source agrees that a label moves evaluation, but it finds the movement stops at the rating: the same label that lowered radiologists' quality scores left their diagnoses unchanged. The split resembles the gap in Can self-ratings replace objective performance scores for AI competence?, where a judgment about the work and the work itself came apart. Here the judgment is advice quality and the work is the diagnosis. The accuracy result also suggests that the ceiling on assisted performance is set by the advice. That concern is raised from another direction by Why does assisted accuracy capture only half the LLM gain?.
The excerpt does not establish how the AI label was assigned or checked, and it omits the methods section and the supplementary by-source results. It covers eight cases drawn from one database, so it says nothing about other imaging tasks or about live clinical workflows. It also contrasts an AI label with a human-expert label on the same advice, which is a different question from how clinicians respond to an AI system that is actually wrong in its own ways. The implication, at the strength the evidence allows, is that a clinical AI tool should be evaluated on what clinicians do with its incorrect outputs, not only on how they rate it. On this evidence, a high quality rating from radiologists would not show that they would catch its errors.
Inquiring lines that read this note 16
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do clinicians calibrate trust in AI medical recommendations?- Can an AI system trained on text consultations handle diagnostic uncertainty in real patient encounters?
- Why do clinicians fail to act on correct AI suggestions in real care?
- What evidence would prove medical AI actually works in clinics?
- Do patients actually perceive AI as worse at addressing their unique medical needs?
- Why did primary care physicians review only 73% of AI-generated transcripts?
- Can clinicians reliably distinguish high-quality AI advice from low-quality advice by appearance alone?
- Do physicians follow incorrect advice more when they trust its source?
- How much diagnostic accuracy is gained when physicians receive expert advice?
- Does this colonoscopy finding apply to other medical specialties using AI?
- Can annotation and explanation labels reduce automation bias in clinical settings?
- Why do nurses misclassify emergencies differently with misleading AI assistance?
- Does AI change clinician cognition or just increase reliance on predictions?
- Do expert physicians also prefer AI-written medical text when it is unlabeled?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can people tell which medical advice is accurate based only on how it reads?
- How does trusting wrong AI advice change what medical action people decide to take?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does the label on advice shape how clinicians judge it?
When clinicians believe advice comes from an expert, do they rate it higher regardless of who actually wrote it? This matters because it reveals whether judgments track the advice itself or just its claimed source.
same kind of label, but there it shifts preference, while here it shifts ratings and not decisions
-
Why does assisted accuracy capture only half the LLM gain?
When an AI system improves on a task, how much of that improvement actually reaches people using it with the AI? Understanding this gap matters because it shows whether complementary strengths automatically translate to better team performance.
the accuracy result implies assisted performance tracks advice correctness, which bounds any synergy
-
Can self-ratings replace objective performance scores for AI competence?
Do people's perceptions of their own AI competence match what they can actually do? This matters because assessment systems might rely on the wrong type of measure to evaluate workplace readiness.
parallel divergence between a judgment and actual performance, here about advice quality
-
How much does wrong AI advice harm radiologist accuracy?
When mammography radiologists receive incorrect AI suggestions labeled as system output, how much does their diagnostic accuracy decline? This matters for understanding automation bias in clinical workflows.
evidence for: wrong BI-RADS advice labeled as AI cut radiologists' accuracy, novices most sharply, showing inaccurate advice lowers accuracy
-
Does time pressure make AI advice more persuasive to experts?
When pathologists work under time constraints, does pressure to decide quickly make them more likely to trust and act on AI recommendations, even when those recommendations are wrong?
evidence for: erroneous AI advice overturned correct estimates in 7% of AI-assisted pathology assessments, showing inaccurate advice degrades accuracy
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Do as AI say: susceptibility in deployment of clinical decision-aids
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Towards Conversational Diagnostic AI
- Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
- Automation Bias in Mammography: The Impact of AI BI-RADS Suggestions on Reader Performance
- GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs
Original note title
radiologists rated advice lower when labeled as AI, but diagnostic accuracy followed whether the advice was correct, not its label