When an AI's medical answer gets a human edit first, do doctors rate it differently than the raw draft?
How much do edited AI responses versus raw outputs affect clinician ratings?
This explores whether clinicians judge AI-written medical or therapeutic responses differently when a human has edited them first than when they see the model's unedited output.
This explores whether a human editing pass changes how clinicians rate AI responses, compared with seeing the model's raw output. The short answer is that the collection has no study that tests this directly. None of the notes compares clinician ratings of edited AI drafts against unedited ones. What the collection does have is evidence about the two sides of that comparison. It suggests that the gap between edited and raw output may be smaller than you'd expect, and that ratings may not measure what we hope they do.
Start with what clinicians actually do when shown AI advice. In one blinded comparison of 104 pairs of responses, clinicians rated GPT-4's advice as more emotionally empathetic than expert advice and found no meaningful difference in scientific quality. They also guessed which response came from the AI only 45% of the time, which is no better than chance Can clinicians tell GPT-4 advice apart from expert advice?. If clinicians can't tell AI text from expert text, the question of whether editing helps becomes harder to answer through ratings alone. The baseline already reads as human-quality.
The writing-assistance research suggests why editing may change less than people assume. Writers edited AI-generated paragraphs only 23% of the time, and their edits left the text about 96% similar to the original Do writers actually edit AI-generated text before publishing?. In practice, "edited" output is often close to raw output. A related study found that AI assistance shifted readers' impressions of the writer on all 29 measured traits, toward more confident, more agreeable and more polished Does AI writing assistance change how readers perceive the writer?. If those qualities survive light editing, a clinician rating an edited response is mostly rating the model's voice.
The less obvious part is that editing out the AI's habits may lower ratings. When researchers trained reward models to remove those persona distortions, writers liked the resulting text less. The clarity and confidence readers enjoy come from the same tendencies that cause the distortions Can AI writing assistance remove distortion without losing appeal?. Evaluators also have a general weakness for style. Models trained to imitate ChatGPT fooled human raters with confident, fluent prose without becoming more factually accurate Can imitating ChatGPT fool evaluators into thinking models improved?. Fluent text also tends to read as competent text Does processing ease mislead users about their own competence?. Together these suggest a risk: careful clinical editing that adds hedges, caveats or blunt corrections could make a response more accurate while lowering its rating.
What the collection can't tell you is the actual size of the effect in clinical settings. To find out, you'd want a study that rates three versions of the same responses (raw, lightly edited and substantially edited) and scores accuracy separately from perceived empathy. Until someone runs that study, the safest reading is that clinician ratings mostly track how a response sounds. Whether a human edited it is a separate question.
Sources 6 notes
Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.
Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.
A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.
Training reward models successfully reduced measured persona distortions, but also reduced writer acceptance of the output. This suggests desirable properties like clarity and confidence operate through the same generative tendencies that produce problematic distortions.
Imitation models fool human evaluators by mimicking ChatGPT's confident, fluent style while failing to improve factuality or generalization on novel tasks. The ceiling is set by base model capability, not fine-tuning method—better fundamentals, not shortcuts, drive real improvement.
Show all 6 sources
High-quality AI output triggers a metacognitive heuristic: users experience fluency as a signal of their own capability, even though they didn't generate it. This self-directed fluency illusion systematically inflates perceived competence because LLMs optimize for fluency regardless of user understanding.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Measuring and Mitigating Persona Distortions from AI Writing Assistance
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- Understanding Reader Perception Shifts upon Disclosure of AI Authorship
- What Influences Readers' and Writers' Perceived Necessity of AI Disclosure?
- "It was 80% me, 20% AI": Seeking Authenticity in Co-Writing with Large Language Models
- Evaluating Large Language Models in Theory of Mind Tasks
- The False Promise of Imitating Proprietary LLMs
- The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows