INQUIRING LINE

Readers rated research abstracts a human wrote and an LLM polished as clearer, even knowing who wrote what, but is that about substance?

Why do readers rate LLM-edited text more favorably?

This explores why people judge text that an LLM has polished, not written from scratch, as better, and whether that preference comes from real quality or from surface features that readers respond to.


This explores why LLM-edited text wins with readers, and whether that preference is about substance or surface. The most direct evidence comes from a study of research abstracts. ML-literate readers could not reliably tell LLM-written abstracts from human ones and tended to assume a human was involved in all of them. Yet the abstracts a human wrote and an LLM then edited got the highest clarity ratings, and readers preferred them 55% of the time even when told who wrote what Can readers tell LLM abstracts from human ones?. So the edge doesn't come from readers spotting 'AI' and favoring it, or from them being fooled. The editing seems to make the prose easier to read, and readers reward that openly.

The corpus doesn't pin down exactly what the editing changes. Studies of other evaluators suggest a strong candidate: presentation moves ratings independently of content. When manuscripts were rewritten to change only the rhetoric (how evidence is framed, how boldly novelty is claimed), AI reviewers' scores shifted even though the science was identical How much does rhetorical style shift AI review scores?. Human readers show a similar pattern. In 24,000 search-chat comparisons, users preferred answers with more citations whether or not those citations were relevant Do users trust citations more when there are simply more of them?. LLM judges fall for fake authority signals and rich formatting Can LLM judges be fooled by fake credentials and formatting?. LLM editing reliably produces fluent, well-organized, confident prose, which is the kind of presentation all these evaluators respond to.

When the reader is itself an LLM, the effect gets stronger and the cause gets clearer. Eight of nine models preferred resumes they had rewritten over matched human versions, and the bias came from matching style, not from better content Do language models favor resumes they rewrote themselves?. LLM judges picked LLM arguments 62% of the time, while human judges split roughly evenly Do LLM judges systematically favor arguments from other LLMs?. One study found that the better a model gets at recognizing its own writing, the more it prefers that writing Do LLMs favor their own text because they recognize it?. Humans don't show this self-recognition loop, but the comparison is telling. It suggests that 'LLM style' is a real, detectable signature, and that at least some evaluators reward familiarity with it rather than quality.

Two results warn against reading too much into this. First, some apparent pro-LLM bias turns out to be an artifact of the data. Across 125,000 peer reviews, LLM-assisted reviewers seemed to favor LLM-assisted papers, but the effect disappeared once paper quality was controlled for Do LLM reviewers actually favor LLM-written papers?. Second, sometimes LLM editing simply makes the text better. At ICLR 2025, reviewers who revised their reviews after LLM feedback produced reviews that blinded raters judged more informative and clearer Can LLM feedback help peer reviewers improve their own reviews?. The unexpected takeaway: readers are partly rewarding genuine clarity and partly rewarding the look of quality, and the corpus doesn't yet have a study that separates the two for human readers.


Sources 9 notes

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

How much does rhetorical style shift AI review scores?

Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.

Do users trust citations more when there are simply more of them?

Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Do language models favor resumes they rewrote themselves?

Across a controlled experiment on 2,245 resumes, eight of nine LLMs preferred their own rewrites over matched human versions when evaluating candidates, with preference rates ranging from 26% to 98%. The bias strengthened in larger models and emerged from stylistic alignment rather than content quality differences.

Show all 9 sources
Do LLM judges systematically favor arguments from other LLMs?

LLM judges selected LLM arguments as winners 62% of the time versus humans' 37%, while humans split votes 39% LLM / 37% human. This same-author bias operates downstream of component scoring and compounds existing judge vulnerabilities, creating a calibration ceiling in RLAIF pipelines.

Do LLMs favor their own text because they recognize it?

Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.

Do LLM reviewers actually favor LLM-written papers?

Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.