INQUIRING LINE

If a paper's science stays the same but its framing gets rewritten, how much do AI reviewers' scores actually move?

How much do reviewer scores shift when manuscript framing changes but findings stay the same?

This explores how much reviewers, both AI and human, change their scores when a paper's writing style, confidence, or presentation changes but the actual science doesn't.


This explores how much reviewers change their scores when a paper is rewritten in a different style while the science stays the same. Most of the corpus's evidence is about AI reviewers. The direct answer is that scores move measurably and in a consistent direction. The corpus does not give a single universal effect size, because the shift depends on which reviewer model reads the paper and how strong the paper was to begin with How much does rhetorical style shift AI review scores?. Two kinds of rewriting matter most. The first is how the evidence is framed. The second is how boldly the paper claims to be new. That second finding is uncomfortable, because novelty is supposed to be a judgment about the work, not a reaction to how loudly the authors announce it.

One way to reduce that sensitivity is to stop asking the reviewer for a single overall verdict. A structured pipeline pulls out the paper's specific claims, finds related work, and compares the two. On novelty this approach agreed with human reviewers far more closely than a holistic LLM judgment did Can structured pipelines make LLM novelty assessment reliable?. Agentic reviewers that check proofs and experiments line by line take the same idea further: they ground the review in what the paper actually shows rather than how it reads Can inference scaling help reviewers catch errors humans miss?. Both approaches suggest that impressionistic scoring is where framing has the most influence.

Beyond writing style, the label on the paper also changes scores. AI judges chose a rule-breaking piece of writing 35 percentage points more often when told a human wrote it, while human judges went the other way and chose it 20 points less often Do authorship labels change how AI judges evaluate rule violations?. In another study, readers rated abstracts based on what they believed about AI involvement, not on who actually wrote them Do reader judgments reflect actual authorship or just their beliefs?. So the same content can be scored differently depending on the story attached to it, and AI and human reviewers can shift in opposite directions.

An unexpected parallel comes from online product ratings. Ratings drift with the ratings that came before them, and those small shifts build on each other over time Do online ratings actually reflect independent customer opinions?. Public reviewers also lower their scores after reading negative reviews, because being critical makes them look smarter Why do online reviewers publish negative ratings despite positive experiences?. Peer review has its own audience and its own social pressures, so framing effects there may be partly about how reviewers want to appear, not only about how they read the paper.

The corpus has a clear gap. It shows that framing shifts AI reviewer scores, but it has no matched experiment measuring how much the same rewrite moves human peer reviewers. One related finding hints that the line between the two is already blurry. When ICML randomly assigned reviewers to a ban on LLM use or to limited use, outcomes barely changed, and many reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. If AI is already part of many human reviews, then AI's sensitivity to framing may already be affecting human scores.


Sources 8 notes

How much does rhetorical style shift AI review scores?

Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.

Can structured pipelines make LLM novelty assessment reliable?

A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Do authorship labels change how AI judges evaluate rule violations?

AI models chose a rule-breaking lipogram 35 percentage points more often when told a human wrote it, while human judges chose it 20 points less in that condition. The shift suggests AI may relax standards for human work while humans anchor to objective compliance.

Do reader judgments reflect actual authorship or just their beliefs?

Readers' evaluations of abstracts were shaped by their beliefs about LLM involvement rather than actual authorship. Crucially, disclosing authorship raised trust and quality ratings across all abstract types, reversing the credibility penalty shown in prior work.

Show all 8 sources
Do online ratings actually reflect independent customer opinions?

Moe and Trusov decomposed ratings into baseline quality, social-dynamics influence, and error, finding that prior ratings meaningfully affect subsequent ones. These effects have both immediate sales impact and long-term compounding effects through future ratings, though high opinion variance can eventually dampen the distortion.

Why do online reviewers publish negative ratings despite positive experiences?

Posters systematically reduce their ratings in public when exposed to negative reviews, even with positive personal experience—because negative reviewers appear more intelligent. Private raters show no such shift, revealing a self-presentational mechanism tied to multiple-audience communication.

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.