SYNTHESIS NOTE
Topics›Domain Specialization›this note

Does labeling advice as AI change how clinicians use it?

When physicians know diagnostic advice comes from an AI system rather than a human expert, do they rely on it differently? This matters because AI labels might trigger skepticism that affects clinical decisions.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

Gaube et al. tested whether telling physicians that diagnostic advice came from "an AI system" changes how they use it. Radiologists (n = 138) and internal or emergency medicine physicians (n = 127) reviewed eight chest X-ray cases. Each case came with advice that was either accurate or inaccurate and labeled as coming from an AI system or an experienced radiologist. All of the advice was written by human experts, so only the label varied. The label moved ratings in one group: "only participants with higher task expertise showed algorithmic aversion by rating the quality of advice to be significantly lower when it came from the AI in comparison to the human." It did not move diagnoses: "the purported source of the advice did not affect participants' performance." What did move accuracy was the advice itself. Task experts performed 40.10% better, and non-experts 37.53% better, when they received accurate rather than inaccurate advice.

The paper's account of the diagnostic result is that advice pulls judgment toward itself. The authors say the advice "could have engaged cognitive biases, by anchoring participants to a particular diagnosis, and triggering confirmatory hypothesis testing," and they describe "a general tendency for participants to agree with advice," stronger among physicians with less task expertise. They define clinical susceptibility as "the propensity to follow incorrect advice." On that measure, 41.73% of IM/EM physicians and 27.54% of radiologists always gave the wrong diagnosis when the advice was wrong. Susceptibility was not limited to weaker performers: 28.26% of radiologists and 17.32% of IM/EM physicians refuted all the incorrect advice they saw. The by-source breakdown of susceptibility sits in a supplementary figure that the excerpt does not include.

This sits against the claim that labels shape reliance. Does the label on advice shape how clinicians judge it? reports a label moving clinicians' preferences. This source agrees that a label moves evaluation, but it finds the movement stops at the rating: the same label that lowered radiologists' quality scores left their diagnoses unchanged. The split resembles the gap in Can self-ratings replace objective performance scores for AI competence?, where a judgment about the work and the work itself came apart. Here the judgment is advice quality and the work is the diagnosis. The accuracy result also suggests that the ceiling on assisted performance is set by the advice. That concern is raised from another direction by Why does assisted accuracy capture only half the LLM gain?.

The excerpt does not establish how the AI label was assigned or checked, and it omits the methods section and the supplementary by-source results. It covers eight cases drawn from one database, so it says nothing about other imaging tasks or about live clinical workflows. It also contrasts an AI label with a human-expert label on the same advice, which is a different question from how clinicians respond to an AI system that is actually wrong in its own ways. The implication, at the strength the evidence allows, is that a clinical AI tool should be evaluated on what clinicians do with its incorrect outputs, not only on how they rate it. On this evidence, a high quality rating from radiologists would not show that they would catch its errors.

Inquiring lines that read this note 16

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do clinicians calibrate trust in AI medical recommendations?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 88 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

radiologists rated advice lower when labeled as AI, but diagnostic accuracy followed whether the advice was correct, not its label