If no one tells doctors which medical answer is AI-written, do they still rate it as highly as a human physician's?
Do expert physicians also prefer AI-written medical text when it is unlabeled?
This explores whether doctors, like lay readers, rate AI-written medical text as highly as (or more highly than) human-written text when they don't know where it came from, and whether their expertise protects them from that effect.
This explores whether doctors, like ordinary readers, rate AI-written medical text as well as or better than human-written text when nobody tells them where it came from. The short answer is that the collection doesn't contain a study that tests exactly this. It has a direct test on patients, a closely related test on radiologists, and enough surrounding evidence to suggest that expertise may matter less than you'd expect.
Start with the non-experts. In a 300-person study, people couldn't tell AI medical answers from doctors' answers any better than chance. They also rated low-accuracy AI answers as trustworthy enough to act on, sometimes more than real doctors' advice Can people tell AI medical advice from doctors' responses?. The same thing shows up outside medicine. Readers gave unlabeled AI-assisted emails the same ratings as human-written ones, and became skeptical only once the AI's role was disclosed Do readers trust unlabeled AI-written messages as much as human ones?. Without a label, people seem to trust by default.
The closest evidence on experts is a radiologist study. Radiologists rated identical advice lower when it was labeled as coming from AI than when it was labeled as coming from a human expert. But their diagnostic accuracy depended only on whether the advice was correct, not on what the label said Does labeling advice as AI change how clinicians use it?. So for clinicians, the label changed how they felt about the advice but not how they used it. That points to a likely answer: take the label away and physicians would probably judge AI text on its content, much as patients do. The difference is that experts may be better at noticing when that content is wrong.
Would expertise help doctors spot AI text in the first place? Probably not. AI-generated text is measurably different from human writing on six measures of vocabulary variety, yet human judges, including trained linguists, can't reliably pick it out. Newer models drift further from human style while becoming harder to detect Can humans detect AI text if machines can measure it?. Peer review shows something similar: a fully AI-generated paper scored above the acceptance threshold in double-blind review at an ICLR workshop Can AI-generated papers pass peer review undetected?. Being an expert in a field doesn't seem to make someone better at telling who wrote a text. And if AI's diagnostic reasoning really does beat physician baselines, as one study reports Can language models reason better than physicians at diagnosis?, unlabeled AI text may often deserve the higher rating.
The more surprising lesson is that the label itself is a design decision with real effects. Disclosing AI involvement lowers ratings only slightly but consistently Does disclosing AI assistance make readers trust articles less?. AI help also changes how readers see the writer, making them seem more confident and polished Does AI writing assistance change how readers perceive the writer?. Writers rarely edit AI drafts before passing them on Do writers actually edit AI-generated text before publishing?. Put together, the question for medicine may be less whether doctors prefer unlabeled AI text and more what the label is for. In the radiology study, the label changed how doctors felt but not what they did, while for patients, the missing label let wrong answers through.
Sources 9 notes
A 300-participant study found participants could not reliably distinguish AI-generated medical responses from doctors' (50% accuracy, chance level) and rated low-accuracy AI answers as valid and trustworthy enough to act on them—comparable to or stronger than their trust in actual doctors' advice.
In a preregistered experiment (N=647), recipients rated unlabeled AI-assisted emails indistinguishably from human-written ones. Only explicit AI disclosure triggered strong skepticism. Recipients appear to default to trust rather than suspicion when origin is unrevealed.
Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.
LLM-generated text differs significantly on six lexical diversity dimensions, confirmed through statistical analysis across multiple models. Yet human judges, including trained linguists, cannot reliably detect these differences—and newer models diverge further while becoming harder to spot.
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
Show all 9 sources
In physician-adjudicated vignette experiments and an emergency room study, a large language model outperformed hundreds of physicians on differential diagnosis, reasoning, triage, and clinical management tasks across multiple touchpoints.
Both human raters (n=1,970) and LLM raters (n=2,520) scored an identical news article lower when it included an AI disclosure statement, but the penalty was small—less than 0.15 points on a 7-point scale.
A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- What Influences Readers' and Writers' Perceived Necessity of AI Disclosure?
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
- Towards Conversational Diagnostic AI
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Measuring and Mitigating Persona Distortions from AI Writing Assistance
- "It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models