Why do people trust the same medical advice more or less just because of the label stuck on it?
Why does labeling advice as AI from a doctor change how people trust it?
This explores why the same medical advice gets trusted differently depending on whether it is labeled as coming from AI or from a doctor, and whether that label actually changes what people do with the advice.
This explores why identical medical advice earns more or less trust depending on whether it carries an AI label or a doctor's name, and whether that label changes anything beyond opinions. The short answer from the corpus is that the label strongly shapes how people *rate* advice, but it barely affects how they *use* it. That makes the label a weaker safeguard than it looks.
Start with what happens when there is no label. Neither patients nor clinicians can reliably tell AI-written medical answers from doctors' answers. In one study, 300 participants guessed at chance level and rated even low-accuracy AI answers as trustworthy enough to act on, sometimes more than real doctors' advice Can people tell AI medical advice from doctors' responses?. Clinicians did no better: blinded experts identified GPT-4 versus expert advice correctly only 45% of the time, and they rated GPT-4 as more emotionally empathetic Can clinicians tell GPT-4 advice apart from expert advice?. Outside medicine the pattern is the same. Unlabeled AI-assisted messages get the same default trust as human ones until someone discloses the AI's role Do readers trust unlabeled AI-written messages as much as human ones?.
So the trust gap comes from the label itself, not from anything in the text. Clinicians preferred advice they *believed* was expert-written 93.55% of the time, even though their beliefs about authorship were no better than coin flips. Their quality and empathy scores followed the perceived author rather than the actual one Does the label on advice shape how clinicians judge it?. The surprise comes when you check behavior. Radiologists rated AI-labeled advice lower, yet their diagnostic accuracy depended only on whether the advice was correct, not on who it was attributed to Does labeling advice as AI change how clinicians use it?. The label changes what people say about the advice. It doesn't seem to change how much the advice pulls their decisions.
That matters because the pull is lopsided. When advice labeled as AI output was wrong, experienced radiologists dropped from 82% to 45.5% accuracy, and inexperienced ones fell below 20% How much does wrong AI advice harm radiologist accuracy?. In an ICU simulation, misleading AI predictions hurt nurse performance nearly twice as much as correct ones helped Do wrong AI predictions hurt more than right ones help?. The same asymmetry shows up in AI fact-checking, where mislabeling true headlines lowered belief in them Does AI fact-checking actually help people spot misinformation?. Skepticism toward the label, then, does not protect people from bad content. Outside medicine the label's penalty can also be small: AI disclosure on a news article cost under 0.15 points on a 7-point scale Does disclosing AI assistance make readers trust articles less?.
The takeaway you may not have expected is that arguing over labels may be arguing over the wrong thing. If the label mostly shapes reputation and the content drives decisions, the more useful designs change how advice is delivered. One example is having AI point out which parts of a case deserve attention rather than handing over a verdict, which reduces anchoring on the machine's answer Can AI guidance reduce anchoring bias better than AI decisions?. One caveat: the evidence on default trust comes from snapshot studies. Whether people grow more suspicious of unlabeled advice as awareness of AI spreads remains an open question Does trust in unlabeled AI messages decline as awareness grows?.
Sources 11 notes
A 300-participant study found participants could not reliably distinguish AI-generated medical responses from doctors' (50% accuracy, chance level) and rated low-accuracy AI answers as valid and trustworthy enough to act on them—comparable to or stronger than their trust in actual doctors' advice.
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
In a preregistered experiment (N=647), recipients rated unlabeled AI-assisted emails indistinguishably from human-written ones. Only explicit AI disclosure triggered strong skepticism. Recipients appear to default to trust rather than suspicion when origin is unrevealed.
Clinicians preferred advice they believed was expert-written 93.55% of the time, even though their guesses about authorship were at chance level. Their scores for quality and empathy shifted based on perceived author, not the text's actual origin.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
Show all 11 sources
A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.
In an ICU simulation, misleading AI predictions degraded nurse performance by 96–120%, while correct predictions improved it by only 53–67%. This asymmetry was hidden by standard metrics that average gains and losses together.
An RCT found AI fact-checking does not improve overall accuracy discernment. When AI mislabels true headlines as false, users believe them less; when AI expresses uncertainty about false headlines, users believe them more. Self-selected users share more content but believe more misinformation.
Both human raters (n=1,970) and LLM raters (n=2,520) scored an identical news article lower when it included an AI disclosure statement, but the penalty was small—less than 0.15 points on a 7-point scale.
Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.
In a single study of 647 participants, readers rated unlabeled AI-assisted messages as favorably as human-written ones. The authors predict awareness may shift this baseline but acknowledge their snapshot design cannot measure whether that erosion actually occurs.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice
- Beyond "Made with AI": Visualizing Provenance Density to Mitigate the Transparency Penalty
- People Defer to AI Moral Advice, But Not Blindly
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- Combining Human Expertise with Artificial Intelligence: Experimental Evidence from Radiology
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content