INQUIRING LINE

AI health advice now sounds just as convincing as a real doctor's — even to doctors, even when it's wrong.

Can people tell which medical advice is accurate based only on how it reads?

This explores whether readers, including doctors, can judge if medical advice is correct just from its wording and tone, with no outside evidence or source information to go on.


This explores whether you can judge medical advice on its wording alone, with no source information or outside evidence to check it against. The corpus gives a fairly blunt answer: mostly no. Readers judge how advice sounds, and AI-generated advice now sounds as good as a doctor's. In a 300-person study, participants guessed whether a response came from an AI or a doctor at chance level. They also rated low-accuracy AI answers as valid and trustworthy enough to act on, sometimes more than real doctors' advice Can people tell AI medical advice from doctors' responses?. The pattern holds outside medicine: when readers had no cues about where claims came from, they fell for fluent made-up claims as readily as true ones Can readers tell truth from fabrication without evidence signals?.

Training doesn't fix this. Clinicians reading GPT-4 and expert advice side by side identified the source at chance (45%), rated GPT-4 as more emotionally empathetic, and found no real difference in scientific quality Can clinicians tell GPT-4 advice apart from expert advice?. Even ML experts couldn't reliably spot LLM-written research abstracts, and they rated LLM-edited ones the clearest Can readers tell LLM abstracts from human ones?. Polish, empathy and clarity are exactly what language models are good at producing, so these qualities say little about whether the content is correct.

There's a twist. When wording can't tell people what to trust, labels take over. Clinicians preferred advice they *believed* was expert-written 93.55% of the time, even though their guesses about who wrote it were no better than chance Does the label on advice shape how clinicians judge it?. But a radiology study found a split between what clinicians say and what they do. They rated advice lower when it was labeled AI, yet their diagnostic accuracy depended on whether the advice was actually correct, not on the label Does labeling advice as AI change how clinicians use it?. So people's stated opinions follow the label, while their decisions follow the advice itself, for better or worse.

That is why accuracy matters more than it seems to. Wrong advice that reads well does more harm than right advice does good. In an ICU simulation, misleading AI predictions hurt nurses' performance by roughly twice as much as correct ones helped, and averaged metrics hid that imbalance Do wrong AI predictions hurt more than right ones help?. AI fact-checkers have the same lopsided effect. When they mislabel true claims, people trust those claims less, and the overall ability to tell true from false doesn't improve Does AI fact-checking actually help people spot misinformation?. Medicine may be especially exposed because medical AI accuracy depends more on getting facts right than on reasoning well Does medical AI need knowledge or reasoning more?. A factual error can sit inside perfectly reasonable-sounding text.

What does seem to work is showing evidence instead of relying on style. When readers could see which claims had been verified, their ability to tell true from false came back strongly Can readers tell truth from fabrication without evidence signals?. One argument explains why wording fails: human expertise includes knowing what your audience needs to hear in order to judge you. AI produces the confident style of expertise without that underlying work, so the fluency itself becomes misleading Can AI replicate the communicative work experts do?. The takeaway is that reading more carefully won't help much. What helps is attaching proof to the claims themselves.


Sources 10 notes

Can people tell AI medical advice from doctors' responses?

A 300-participant study found participants could not reliably distinguish AI-generated medical responses from doctors' (50% accuracy, chance level) and rated low-accuracy AI answers as valid and trustworthy enough to act on them—comparable to or stronger than their trust in actual doctors' advice.

Can readers tell truth from fabrication without evidence signals?

In an 81-person study, participants given no provenance cues showed no significant truth discernment (p = .43), falling for fluent hallucinations as readily as ground truth. An idealized Provenance Density interface showing verified claims restored a +4.15 point gap (p < .001).

Can clinicians tell GPT-4 advice apart from expert advice?

Blinded clinician ratings of 104 response pairs found GPT-4 advice favored on emotional empathy, with no significant differences in scientific quality or cognitive empathy. Clinicians identified the source at chance level (45% accuracy), suggesting the two were indistinguishable in written form.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Does the label on advice shape how clinicians judge it?

Clinicians preferred advice they believed was expert-written 93.55% of the time, even though their guesses about authorship were at chance level. Their scores for quality and empathy shifted based on perceived author, not the text's actual origin.

Show all 10 sources
Does labeling advice as AI change how clinicians use it?

Radiologists rated AI-labeled advice lower than identical advice labeled human-expert, yet their diagnostic accuracy depended on whether the advice was correct, not its source. This suggests labels shape what clinicians think about advice but not how they use it.

Do wrong AI predictions hurt more than right ones help?

In an ICU simulation, misleading AI predictions degraded nurse performance by 96–120%, while correct predictions improved it by only 53–67%. This asymmetry was hidden by standard metrics that average gains and losses together.

Does AI fact-checking actually help people spot misinformation?

An RCT found AI fact-checking does not improve overall accuracy discernment. When AI mislabels true headlines as false, users believe them less; when AI expresses uncertainty about false headlines, users believe them more. Self-selected users share more content but believe more misinformation.

Does medical AI need knowledge or reasoning more?

The KI/InfoGain framework reveals that medical domain accuracy correlates more strongly with knowledge correctness than reasoning quality, while mathematical domains show the inverse pattern. This distinction has direct implications for which training strategies to prioritize in each domain.

Can AI replicate the communicative work experts do?

Expertise requires anticipating audience acceptability and social validity, not just retrieving information. AI lacks the mechanism to perform this communicative work, making its fluent output epistemically misleading despite its confident form.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.