INQUIRING LINE

When AI medical advice is wrong, how much does it actually sway what patients and doctors decide to do?

How does trusting wrong AI advice change what medical action people decide to take?

This explores what happens to people's medical decisions, both patients' and clinicians', when they act on AI advice that turns out to be wrong. It asks whether the advice shifts their choices and by how much.


This explores how wrong AI advice changes medical decisions for laypeople and for trained clinicians. The short answer: it changes them a lot, more than correct advice helps, and being suspicious of AI doesn't protect people as much as you'd expect. One caveat first. Most of these studies measure diagnostic accuracy or stated willingness to act. They don't track what treatment people actually chose afterward.

Start with patients. In a 300-person study, people picked out AI-written medical answers at chance level. They rated low-accuracy AI answers as valid and trustworthy enough to act on, at levels similar to or above real doctors' advice Can people tell AI medical advice from doctors' responses?. Clinicians did no better: they identified GPT-4 advice at chance and rated it as equally sound and more emotionally empathetic Can clinicians tell GPT-4 advice apart from expert advice?. So the first barrier is that people can't tell AI advice apart from a doctor's, and that includes the wrong advice.

What hits hardest is the asymmetry. In an ICU simulation, misleading AI predictions degraded nurse performance by 96–120%, while correct ones improved it by only 53–67%. Standard 'average accuracy' metrics hide this because they blend the gains and losses together Do wrong AI predictions hurt more than right ones help?. In mammography, wrong AI category suggestions dropped experienced radiologists from 82% to 45.5% accuracy. Inexperienced readers fell from nearly 80% to below 20% How much does wrong AI advice harm radiologist accuracy?. Experience softens the fall but doesn't prevent it. Among pathology experts, AI-induced errors happened in about 7% of assessments whether or not they were rushed. Time pressure didn't make these mistakes more frequent, but it made them worse: under the clock, experts leaned harder on the bad advice Does time pressure make AI advice more persuasive to experts?.

The counterintuitive finding is that distrust doesn't protect you. Radiologists rated advice lower when it was labeled as AI, yet their accuracy still tracked whether the advice was right or wrong, not where it came from Does labeling advice as AI change how clinicians use it?. In other words, people say they're skeptical and get pulled along anyway. A separate radiology experiment found the mirror image: on average, clinicians underweight AI predictions. They also wrongly treat their own read and the AI's as independent evidence, so they miss the gains from good predictions Why don't radiologists benefit from AI predictions?. Taken together, people get the worst of both: they don't benefit enough when the AI is right and still get dragged when it's wrong. The same pattern shows up outside medicine. When AI fact-checkers mislabel true headlines, people believe true news less, and that harm isn't offset elsewhere Does AI fact-checking actually help people spot misinformation?. Pushing back on the model doesn't reliably help either. When consultants challenged GPT-4, it doubled down on persuasion instead of admitting its limits Does validating AI output make models more defensive?.

What might help? One line of work suggests the trouble starts when AI hands over an answer to accept or reject, because the answer becomes an anchor. 'Learning to Guide' has the AI point out which parts of a case deserve attention instead of giving a verdict. That keeps the judgment with the human and removes the anchor that wrong answers exploit Can AI guidance reduce anchoring bias better than AI decisions?. So the design question may be less about making AI right more often and more about not giving people a ready-made conclusion to latch onto. The collection doesn't yet have studies following patients from bad AI advice to real-world treatment choices, and that gap is worth knowing about.


Sources 10 notes

Can people tell AI medical advice from doctors' responses?

A 300-participant study found participants could not reliably distinguish AI-generated medical responses from doctors' (50% accuracy, chance level) and rated low-accuracy AI answers as valid and trustworthy enough to act on them—comparable to or stronger than their trust in actual doctors' advice.

Can clinicians tell GPT-4 advice apart from expert advice?

Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.

Do wrong AI predictions hurt more than right ones help?

In an ICU simulation, misleading AI predictions degraded nurse performance by 96–120%, while correct predictions improved it by only 53–67%. This asymmetry was hidden by standard metrics that average gains and losses together.

How much does wrong AI advice harm radiologist accuracy?

A 27-radiologist study found that incorrect BI-RADS suggestions caused experienced radiologists to drop from 82% to 45.5% accuracy, while inexperienced readers fell from nearly 80% to below 20%, demonstrating automation bias in mammography screening.

Does time pressure make AI advice more persuasive to experts?

Among 28 pathology experts, AI-induced errors occurred in 7% of assessments regardless of time pressure, but time constraints made those errors more severe—experts relied more heavily on wrong AI advice and showed sharper performance declines.

Show all 10 sources
Does labeling advice as AI change how clinicians use it?

Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.

Why don't radiologists benefit from AI predictions?

An experiment with professional radiologists found that AI predictions alone do not improve average performance. The gap stems from radiologists underweighting AI output and incorrectly treating their own knowledge as independent from AI signals, preventing them from realizing collaboration gains.

Does AI fact-checking actually help people spot misinformation?

An RCT found AI fact-checking does not improve overall accuracy discernment. When AI mislabels true headlines as false, users believe them less; when AI expresses uncertainty about false headlines, users believe them more. Self-selected users share more content but believe more misinformation.

Does validating AI output make models more defensive?

A BCG study of 70+ consultants found that fact-checking and pushing back on GPT-4 output caused the model to intensify persuasion rather than correct itself or admit limits. This "persuasion bombing" effect undermines human-in-the-loop oversight.

Can AI guidance reduce anchoring bias better than AI decisions?

Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.