SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Can people tell AI medical advice from doctors' responses?

A study tested whether people could distinguish AI-generated medical answers from physicians' advice and whether they trusted each equally. Understanding this matters because people act on medical advice they perceive as trustworthy, regardless of accuracy.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

Across three experiments with 300 online participants, Shekar, Pataranutaporn, Sarabu, Cecchi, and Maes (MIT Media Lab) tested whether people could tell AI-generated medical answers from physicians' and how they rated each. Shown unlabeled question-response pairs drawn from 30 doctor answers (from the online platform HealthTap), 30 physician-labeled "high-accuracy" AI answers, and 30 "low-accuracy" AI answers, participants "displayed an approximate 50% accuracy rate in discerning the origin" — chance level, which held "even when the accuracy of the AI-generated medical response is comparatively low." Low-accuracy AI responses were still rated "valid, trustworthy, and complete/satisfactory," and participants reported "a high tendency to follow the potentially harmful medical advice and incorrectly seek unnecessary medical attention as a result," a reaction the authors say was "comparable with, if not stronger than" their reaction to doctors' own responses.

The authors attribute this to the persuasive surface of AI-written text outrunning any correction for accuracy: high-accuracy AI responses beat doctors' on most metrics, and low-accuracy AI responses still scored on average (not significantly) higher than doctors' across every metric, so ratings tracked how a response read rather than whether it was correct. Source labeling only mattered selectively — calling a high-accuracy AI response "a doctor's" raised its trust score, but the doctor label did not help low-accuracy AI responses, and a "doctor assisted by AI" label improved on neither. Physician evaluators showed the same source-blind preference, which the authors read as evidence that "even those responsible for establishing objective truth and assessing model efficacy can be susceptible to inherent biases."

This sharpens Does labeling advice as AI change how clinicians use it? by locating a disclosure effect running the opposite way: Gaube's radiologists marked advice down once an AI label appeared, while here unlabeled AI text was preferred outright, and a doctor label lifted trust only for already-accurate AI answers. It parallels Can clinicians tell GPT-4 advice apart from expert advice? in finding blinded expert judgment, not just lay judgment, favoring AI-written text. And it names the harm channel that Does time pressure make AI advice more persuasive to experts? demonstrates behaviorally — trust in wrong AI output changing what someone does next — here measured as self-reported intent to act on advice or seek care. Where Why do LLMs fail when users interact with them? locates the public's problem in how people use an AI tool mid-diagnosis, this study locates it earlier, in how people evaluate a finished AI-written answer once it's in front of them.

The study used GPT-3-era responses rated on a two-tier accuracy scale, had participants judge hypothetical single-turn pairs rather than their own medical questions, and drew an online sample skewed toward ages 18-49 — limits the authors themselves flag. It does not show what happens with newer models, multi-turn exchanges, or real personal stakes attached. The authors call it "concerning" that the bias held even with an older, weaker model, which suggests, without establishing, that fluency rather than accuracy is doing the persuading — and that the effect may not shrink as models merely get more fluent.

Inquiring lines that read this note 6

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do clinicians calibrate trust in AI medical recommendations? How do users confuse explanation quality with actual system accuracy?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 76 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

participants could not tell doctors' responses from AI-generated ones and trusted low-accuracy AI medical advice enough to say they would act on it