Can people tell AI medical advice from doctors' responses?
A study tested whether people could distinguish AI-generated medical answers from physicians' advice and whether they trusted each equally. Understanding this matters because people act on medical advice they perceive as trustworthy, regardless of accuracy.
Across three experiments with 300 online participants, Shekar, Pataranutaporn, Sarabu, Cecchi, and Maes (MIT Media Lab) tested whether people could tell AI-generated medical answers from physicians' and how they rated each. Shown unlabeled question-response pairs drawn from 30 doctor answers (from the online platform HealthTap), 30 physician-labeled "high-accuracy" AI answers, and 30 "low-accuracy" AI answers, participants "displayed an approximate 50% accuracy rate in discerning the origin" — chance level, which held "even when the accuracy of the AI-generated medical response is comparatively low." Low-accuracy AI responses were still rated "valid, trustworthy, and complete/satisfactory," and participants reported "a high tendency to follow the potentially harmful medical advice and incorrectly seek unnecessary medical attention as a result," a reaction the authors say was "comparable with, if not stronger than" their reaction to doctors' own responses.
The authors attribute this to the persuasive surface of AI-written text outrunning any correction for accuracy: high-accuracy AI responses beat doctors' on most metrics, and low-accuracy AI responses still scored on average (not significantly) higher than doctors' across every metric, so ratings tracked how a response read rather than whether it was correct. Source labeling only mattered selectively — calling a high-accuracy AI response "a doctor's" raised its trust score, but the doctor label did not help low-accuracy AI responses, and a "doctor assisted by AI" label improved on neither. Physician evaluators showed the same source-blind preference, which the authors read as evidence that "even those responsible for establishing objective truth and assessing model efficacy can be susceptible to inherent biases."
This sharpens Does labeling advice as AI change how clinicians use it? by locating a disclosure effect running the opposite way: Gaube's radiologists marked advice down once an AI label appeared, while here unlabeled AI text was preferred outright, and a doctor label lifted trust only for already-accurate AI answers. It parallels Can clinicians tell GPT-4 advice apart from expert advice? in finding blinded expert judgment, not just lay judgment, favoring AI-written text. And it names the harm channel that Does time pressure make AI advice more persuasive to experts? demonstrates behaviorally — trust in wrong AI output changing what someone does next — here measured as self-reported intent to act on advice or seek care. Where Why do LLMs fail when users interact with them? locates the public's problem in how people use an AI tool mid-diagnosis, this study locates it earlier, in how people evaluate a finished AI-written answer once it's in front of them.
The study used GPT-3-era responses rated on a two-tier accuracy scale, had participants judge hypothetical single-turn pairs rather than their own medical questions, and drew an online sample skewed toward ages 18-49 — limits the authors themselves flag. It does not show what happens with newer models, multi-turn exchanges, or real personal stakes attached. The authors call it "concerning" that the bias held even with an older, weaker model, which suggests, without establishing, that fluency rather than accuracy is doing the persuading — and that the effect may not shrink as models merely get more fluent.
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do clinicians calibrate trust in AI medical recommendations?- Why did lay users with AI models fail to match unaided physicians on diagnosis?
- Do expert physicians also prefer AI-written medical text when it is unlabeled?
- Why does labeling advice as AI from a doctor change how people trust it?
- Can people tell which medical advice is accurate based only on how it reads?
- How does trusting wrong AI advice change what medical action people decide to take?
Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does labeling advice as AI change how clinicians use it?
When physicians know diagnostic advice comes from an AI system rather than a human expert, do they rely on it differently? This matters because AI labels might trigger skepticism that affects clinical decisions.
opposite disclosure effect: AI labels suppressed trust there, while unlabeled AI text was preferred outright here
-
Can clinicians tell GPT-4 advice apart from expert advice?
This study explores whether trained clinicians can distinguish AI-generated psychological advice from expert advice, and how they rate the quality and empathy of each. The question matters for understanding whether AI might reliably supplement human expertise in mental health settings.
same pattern of blinded expert judgment favoring AI-written text over human expert text
-
Does time pressure make AI advice more persuasive to experts?
When pathologists work under time constraints, does pressure to decide quickly make them more likely to trust and act on AI recommendations, even when those recommendations are wrong?
shows the same overtrust mechanism changing actual professional judgments, not just survey ratings
-
Why do LLMs fail when users interact with them?
Standard benchmarks show LLMs excel at medical diagnosis alone, yet real users get no benefit. This explores where the breakdown happens between model capability and human decision-making.
locates a related public-facing AI-medicine risk earlier, in how people use a tool rather than how they rate a finished answer
-
How much does wrong AI advice harm radiologist accuracy?
When mammography radiologists receive incorrect AI suggestions labeled as system output, how much does their diagnostic accuracy decline? This matters for understanding automation bias in clinical workflows.
Extends the trust risk to experts: mislabeled AI categories cut even experienced radiologists' accuracy nearly in half
-
Do wrong AI predictions hurt more than right ones help?
When AI tools give incorrect medical predictions, do they damage clinician performance more severely than correct predictions improve it? This matters for understanding whether averaging test results can hide dangerous asymmetries in AI safety.
Evidence for the risk: misleading AI predictions hurt nurses' performance far more than correct predictions helped them
-
Does the label on advice shape how clinicians judge it?
When clinicians believe advice comes from an expert, do they rate it higher regardless of who actually wrote it? This matters because it reveals whether judgments track the advice itself or just its claimed source.
Extends the indistinguishability finding: clinicians guessed authorship at chance yet preferred advice believed expert-written 93.55% of the time
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Do as AI say: susceptibility in deployment of clinical decision-aids
- Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice
- Towards Conversational Diagnostic AI
- Clinical knowledge in LLMs does not translate to human interactions
- People Defer to AI Moral Advice, But Not Blindly
- Epistemic Deference to AI
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
Original note title
participants could not tell doctors' responses from AI-generated ones and trusted low-accuracy AI medical advice enough to say they would act on it