INQUIRING LINE

When an AI grades a research paper, does polished, confident wording count for more than the science itself?

Do AI reviews depend more on writing style than scientific merit?

This explores whether AI systems that review research papers can be swayed by how a paper is written (its framing, confidence and polish) instead of by whether the science holds up.


This explores whether AI reviewers judge the writing more than the research. The short answer from the corpus is that writing matters more than it should. In one set of experiments, researchers rewrote papers to change only the rhetoric and left the scientific content the same. LLM reviewer scores still moved in a consistent direction (How much does rhetorical style shift AI review scores?). Two kinds of change moved scores the most: how the evidence was framed, and how boldly the paper claimed to be new. The effect depended on which model was reviewing and on how strong the paper was to begin with. So AI reviewers do pay attention to substance, but presentation shifts their scores in ways a careful reader should not allow.

The bigger problem shows up at scale. A separate study found that simple automatic rewrites of a paper's text raised AI review scores by about 0.45 points with no change to the science (Can AI systems safely replace human peer reviewers?). The same study found a 'hivemind' effect: different AI reviewers agree with each other more than human reviewers do. That combination is the danger. One human reviewer swayed by polish is noise. Many AI reviewers swayed in the same direction is a weakness that authors can exploit, and a rewrite that works on one model will tend to work on the others.

Humans are not immune to this either. A position paper on AI conference reviewing found that human ratings correlate with review length, and it argues that authors, reviewers and venues all share the blame (Can two-stage review and badges fix AI conference peer review?). The AI Scientist-v2 papers make the same point from the other side. One fully AI-generated manuscript scored an average of 6.33 from ICLR workshop reviewers. Its own authors later judged that it fell short of main-conference rigor, and they found a citation error in it after the fact (Can AI systems generate research papers that pass peer review?, Can AI-generated papers pass peer review undetected?). Fluent, well-structured writing can get past reviewers, human or machine, faster than the science behind it can be checked.

There is also a possible feedback loop, though it is an inference and no study in the corpus tests it directly. In a study of 2,939 writers, AI writing help made authors come across as more confident and higher quality on every one of the 29 traits measured (Does AI writing assistance change how readers perceive the writer?). Writers changed the AI's text only 23% of the time, and their edits left it about 96% the same (Do writers actually edit AI-generated text before publishing?). If AI reviewers reward confident framing and AI writing tools add confident framing, papers polished by AI could score better with AI reviewers for reasons unrelated to their science.

The more hopeful finding is that the bias seems to come from how the reviewer is set up, not from AI as such. Reviewers that actually check the work do much better than ones that just read and score. PAT is an agentic reviewer that uses extra compute to go through proofs and experiments line by line. It found real flaws in papers accepted at STOC and ICML that human reviewers had missed (Can inference scaling help reviewers catch errors humans miss?). In a related result, agent-based judges that gather evidence before deciding had about 100 times less 'judge shift' than plain LLM judges (Can agents evaluate AI outputs more reliably than language models?). A third option is to have AI coach human reviewers instead of replacing them. At ICLR 2025, AI feedback on reviews led 27% of reviewers to revise, and blinded raters judged the revised reviews more informative (Can LLM feedback help peer reviewers improve their own reviews?). The useful question is less whether AI reviewers are fooled by style and more whether a given reviewer is set up to check the claims or only to read them.


Sources 10 notes

How much does rhetorical style shift AI review scores?

Rewriting manuscripts' rhetoric while preserving scientific content moves LLM reviewer scores measurably. Evidence framing and novelty stance produce the largest contrasts; effects vary by reviewer model and interact with the paper's original quality tier.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Show all 10 sources
Does AI writing assistance change how readers perceive the writer?

A study of 2,939 writers and 11,091 readers found AI assistance shifted every tested dimension—29 total—toward extremism, confidence, quality, agreeableness, and perceived privilege. Distortions were statistically significant and directional, not random noise.

Do writers actually edit AI-generated text before publishing?

Writers edited AI-generated paragraphs only 23% of the time, with edits averaging 96% similarity to the original. This means AI's opinionated and distorted voice propagates with minimal human filtering before publication.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can agents evaluate AI outputs more reliably than language models?

Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.