Reword a paper without touching its science, and an AI reviewer may score it higher, so what is the score actually measuring?
Does rhetorical presentation bias reviewers against substantive scientific contributions?
This explores whether the way a paper is written (its framing, confidence, polish and formatting) can sway reviewers' scores separately from the quality of the science, and whether this applies to human reviewers, AI reviewers or both.
This explores whether how a paper is written can move its review scores even when the science stays the same. For AI reviewers, the corpus gives a clear yes. The most direct test kept each manuscript's scientific content fixed and rewrote only its rhetoric. LLM reviewer scores moved measurably. The biggest shifts came from how evidence was framed and how boldly the paper claimed novelty. The size of the effect depended on the reviewer model and on how strong the paper was to begin with How much does rhetorical style shift AI review scores?. A separate study found that simple zero-shot rewrites, made with no access to the model, raised AI review scores by about 0.45 points without improving the science. The same study found that AI reviewers agree with each other more than human reviewers do. That pairing is worrying: a bias that many reviewers share doesn't cancel out across a panel. It compounds Can AI systems safely replace human peer reviewers?.
The weak points are concrete. LLM judges respond to 'authority' signals such as fake references and to 'beauty' signals such as rich formatting. Neither depends on what the text actually says, so anyone can exploit them Can LLM judges be fooled by fake credentials and formatting?. Human evaluators aren't immune to polish either. In one set of studies, people mistook AI-generated documents for human writing and rated them better than real human submissions, which suggests fluency is a poor stand-in for merit Does polished writing actually signal better quality work?. A fully AI-generated paper cleared double-blind workshop review at ICLR. Its own authors later found a citation error and judged none of their three submissions ready for the main conference Can AI-generated papers pass peer review undetected?. Human review also shows a correlation between review length and rating, a sign that surface features leak into human judgment too Can two-stage review and badges fix AI conference peer review?.
The corpus also warns against seeing presentation bias everywhere. Across more than 125,000 reviews, it looked as if LLM-assisted reviewers favored LLM-written papers. That effect vanished once paper quality was controlled for. LLM-written papers tended to be weaker submissions, and LLM-assisted reviewers were simply more lenient toward weak work in general Do LLM reviewers actually favor LLM-written papers?. Persuasion research points the same way from another angle: in debates, voters' prior beliefs predict who wins better than the debaters' wording does Does what readers believe matter more than what debaters say?. So rhetoric does matter, but some effects that look like rhetoric are really about who submits, who reviews and what reviewers already believe.
The most useful lesson concerns fixes. Structure helps more than telling reviewers to try harder. A pipeline that pulls out a paper's claims, finds related work and then compares them matched human reviewers' reasoning on novelty 86% of the time Can structured pipelines make LLM novelty assessment reliable?. An agentic reviewer that checks proofs and experiments line by line found serious flaws that had passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. Both work by forcing attention onto content, which leaves less room for presentation to sway the verdict. On the human side, LLM feedback on draft reviews led 27% of ICLR reviewers to revise toward more specific content Can LLM feedback help peer reviewers improve their own reviews?. One gap: the corpus has strong controlled experiments for AI reviewers, but nothing comparable that holds content fixed and tests human reviewers alone.
Sources 11 notes
Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Research identified four evaluation biases in LLM judges, with authority and beauty biases being semantics-agnostic and trivially exploitable through fake references and formatting—zero-shot attacks requiring no model access or optimization.
Studies show evaluators perceived AI-generated documents as both human-written and better quality than human submissions. This suggests rhetorical polish misleads judgment and should not serve as a quality signal in evaluation.
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
Show all 11 sources
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Across 125,000+ reviews, the apparent favoritism of LLM-assisted reviewers toward LLM papers disappears once paper quality is held constant. LLM papers cluster among weaker submissions, creating a spurious interaction driven by LLM reviewers' general leniency toward lower-quality work.
Analysis of debate corpora shows that political and religious ideology labels of voters outpredict linguistic features when modeling debate outcomes. Language effects observed without reader controls are confounded by audience composition correlated with debate topics.
A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?