INQUIRING LINE

Label a text as AI-written and readers trust it less, even when not a single word has changed.

Why does disclosure of LLM authorship change reader trust and preference?

This explores why telling readers that a text was written with an LLM changes how much they trust it and whether they prefer it, even when the text itself is identical.


This explores why labeling a text as AI-written shifts how readers judge it, even though the words on the page stay the same. The corpus's clearest answer is that disclosure doesn't change what readers think of the writing's quality as much as what they think of the writer's intent. In a study of 261 readers, revealing AI authorship lowered perceived trustworthiness, caring, and likability. The drop was steepest in interpersonal writing, because readers saw AI as unable to feel real empathy and read its use as a breach of social expectations How does revealing AI authorship change reader trust?. So the penalty is largest where the reader expected the writer to have personally shown up.

What may surprise you is that before disclosure, most readers can't tell the difference. Even readers with machine learning expertise couldn't reliably pick out LLM-written research abstracts and tended to assume a human was involved. LLM-edited abstracts actually got the highest clarity ratings, and they were still preferred 55% of the time when authorship was disclosed Can readers tell LLM abstracts from human ones?. So disclosure doesn't uncover a flaw readers had already noticed. It adds new information about how the text was made, and how much that information matters depends on the genre. In technical writing where clarity is the point, the penalty can be small. In writing where caring is the point, it is large.

Readers also care about disclosure more than writers expect. In a 727-person vignette study, readers rated disclosure as more necessary than writers did across the board. It felt most necessary when AI text was pasted in directly and couldn't easily be replaced, and how much effort the writer put in made no difference Do readers and writers differ on AI disclosure necessity?. That suggests readers are tracking whose words these are, not how hard someone worked. Human raters applied the disclosure penalty evenly regardless of the author's demographics. LLM raters did something different: they favored Black or women authors when AI use was hidden, and those preferences disappeared once it was disclosed Do LLM raters show hidden demographic preferences that disclosure erases?.

A lateral thread helps explain why a label can carry this much weight: trust often runs on surface signals that aren't tied to content. Users preferred search answers with more citations almost as much when the citations were irrelevant as when they were relevant Do users trust citations more when there are simply more of them?. LLM judges also reward fake references and rich formatting regardless of quality Can LLM judges be fooled by fake credentials and formatting?. An authorship label is another signal of this kind, but it pushes trust down instead of up. Machine judges reverse the human pattern in another way too: LLM judges picked LLM-written arguments as winners 62% of the time, while human voters split roughly evenly Do LLM judges systematically favor arguments from other LLMs?.

The corpus documents the effect well but says less about its mechanism. The 'social expectation violation' explanation rests mainly on one study. No note here tests whether the penalty fades as AI-assisted writing becomes normal, or whether partial disclosure ('edited with AI' vs. 'written by AI') changes the result. Those are open doorways rather than settled answers.


Sources 7 notes

How does revealing AI authorship change reader trust?

A study of 261 readers found that disclosing AI authorship consistently lowered perceived trustworthiness, caring, and likability, with the steepest drops in interpersonal writing like personal interaction. Readers saw AI as incapable of genuine empathy, viewing its use as a violation of social expectations.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Do readers and writers differ on AI disclosure necessity?

A 727-person vignette study found readers consistently rated AI disclosure as more necessary than writers did. Disclosure seemed most necessary when AI text was directly incorporated and irreplaceable, while writer effort had no effect on these judgments.

Do LLM raters show hidden demographic preferences that disclosure erases?

GPT-4o-mini showed pronounced preference for Black authors and Qwen2.5-7B-Instruct favored women authors when AI use was undisclosed, but both preferences vanished under disclosure. Human raters showed uniform disclosure penalties regardless of author demographics.

Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Show all 7 sources
Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Do LLM judges systematically favor arguments from other LLMs?

LLM judges selected LLM arguments as winners 62% of the time versus humans' 37%, while humans split votes 39% LLM / 37% human. This same-author bias operates downstream of component scoring and compounds existing judge vulnerabilities, creating a calibration ceiling in RLAIF pipelines.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.