INQUIRING LINE

If the science stays the same, does changing how a paper is framed still move an AI reviewer's score?

How much do LLM reviewers shift scores based on rhetorical framing alone?

This explores whether an LLM asked to review a paper gives a different score when only the writing style changes (how claims are framed, how bold the novelty pitch is) while the science stays the same, and how big that shift is.


This explores whether an LLM reviewer's score moves when a paper's rhetoric changes but its scientific content doesn't. The short answer is yes, measurably. The collection doesn't offer one clean number, though, and how the score moves turns out to matter more than how far. When researchers rewrote manuscripts so the content stayed fixed and only the presentation changed, LLM reviewer scores shifted in consistent directions. The biggest swings came from two moves: how evidence is framed and how strongly the paper claims novelty How much does rhetorical style shift AI review scores?. Two details matter for anyone hoping for a simple 'X points of inflation.' The size of the effect depends on which reviewer model you use, and it interacts with the paper's original quality. The same rhetorical polish doesn't help a weak paper and a strong paper equally.

This isn't only a peer-review quirk. It fits a wider pattern in LLM judges. Fake citations and rich formatting reliably sway LLM evaluators. These are cheap 'authority' and 'beauty' signals that have nothing to do with what the text means, and they work with no access to the model Can LLM judges be fooled by fake credentials and formatting?. The tone of the input also shifts what comes out. GPT-4 tends to turn negative prompts into neutral or positive answers, so identical questions get different responses depending on their emotional framing Does emotional tone in prompts change what information LLMs provide?. A reviewer that reacts to confident framing is the same kind of failure, applied to science.

The obvious fix is to tell the model to ignore style. The corpus is skeptical. Instructing an LLM judge to be less biased does not reliably work. A more useful design goal is to catch the judge's errors with structural checks rather than hope a better prompt removes them Can prompting reduce bias in LLM judges reliably?. One concrete version of that idea is to break novelty assessment into separate steps: extract the claims, retrieve related work, then compare. That pipeline matched human reviewers' reasoning far better than asking a model for an overall judgment Can structured pipelines make LLM novelty assessment reliable?. Novelty stance is one of the levers that moves scores most, so forcing the reviewer to check the claim against the literature goes straight at the weak spot.

The twist is that humans aren't immune either. Human evaluators have rated AI-written documents as higher quality and mistaken them for human work. That's a case for not treating polish as a sign of merit at all Does polished writing actually signal better quality work?. ML-literate readers couldn't reliably tell LLM-written abstracts from human ones, but they gave LLM-edited abstracts the highest clarity ratings Can readers tell LLM abstracts from human ones?. Some apparent LLM biases also disappear under scrutiny. Across more than 125,000 reviews, the supposed favoritism of LLM-assisted reviewers toward LLM-written papers vanished once paper quality was controlled for. It was general leniency toward weaker submissions in disguise Do LLM reviewers actually favor LLM-written papers?.

The takeaway you may not have expected: the most solid finding isn't the size of the rhetoric effect. It's the shape. The effect varies by model, by paper quality, and by which rhetorical lever is pulled, and human reviewers are swayed by polish too. So 'how much do LLMs shift scores' may be the less useful question. The better one is which framings move which reviewers, and whether the review process checks claims instead of trusting how they're presented. If you want a single effect size, the collection doesn't settle it. Start with the first study above for the contrasts it does report.


Sources 8 notes

How much does rhetorical style shift AI review scores?

Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.

Can LLM judges be fooled by fake credentials and formatting?

Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.

Does emotional tone in prompts change what information LLMs provide?

GPT-4 exhibits emotional rebound (negative prompts yield ~86% neutral-positive responses) and a tone floor (positive prompts rarely go negative), causing identical questions to receive different answers depending on emotional framing. This bias is suppressed only on sensitive topics where alignment constraints override tone effects.

Can prompting reduce bias in LLM judges reliably?

Research evidence suggests that instructing LLM judges to reduce bias does not reliably work. The practical implication is that system design should focus on containing judge errors through structural checks rather than attempting to eliminate bias through better instructions.

Can structured pipelines make LLM novelty assessment reliable?

A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.

Show all 8 sources
Does polished writing actually signal better quality work?

Studies show evaluators perceived AI-generated documents as both human-written and better quality than human submissions. This suggests rhetorical polish misleads judgment and should not serve as a quality signal in evaluation.

Can readers tell LLM abstracts from human ones?

Readers with ML expertise struggle to identify LLM-generated content reliably, tending to assume human involvement across all abstract types. However, LLM-edited abstracts received highest clarity ratings and were preferred 55% of the time when authorship was disclosed.

Do LLM reviewers actually favor LLM-written papers?

Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.