Stop Automating Peer Review Without Rigorous Evaluation

Paper · arXiv 2605.03202 · Published May 4, 2026
Domain Specialization in LLMs

Large language models offer a tempting solution to address the peer review crisis. This position paper argues that today’s AI systems should not be used to produce paper reviews. We ground this position in an empirical comparison of humanversus AI-generated ICLR 2026 reviews and an evaluation of the effect of automated paper rewriting on different AI reviewers. We identify two critical issues: 1) AI reviewers exhibit a hivemind effect of excessive agreement within and across papers that reduces perspective diversity. 2) AI review scores are trivially gameable through paper laundering: prompting an LLM to rewrite a paper could significantly increase the scores from AI reviewers, demonstrating that LLM reviewers are easy to game through stylistic changes rather than scientific results. However, non-gameability and review diversity are necessary but not sufficient conditions for automation. We argue that addressing the peer review crisis requires a science of peer review automation—not general-purpose LLMs deployed without rigorous evaluation.1

Introduction. Scientific peer review is the guarantor for scientific discovery and credibility. However, it faces many challenges (Shah, 2022; Lin et al., 2025a): Submission volumes grow faster than reviewer pools can expand, and LLMwritten reviews are steadily increasing (Liang et al., 2024a; Russo et al., 2025; Emi, 2025). Conference organizers, seeking to deliver timely decisions, have begun automating parts of the process. AAAI 2025 trialed LLM-generated reviews alongside human reviews (AAAI, 2025). Some venues now experiment with fully automated AI reviewer agents (Bianchi et al., 2025b). This trajectory raises a critical question: which parts of peer review, if any, should be automated? In this paper, we argue that answering this question needs new tools and rigorous empirical evaluation.

Having a paper accepted at a top AI conference can change someone’s career. This impact makes automated peer review a high-stakes AI application, and those applications demand rigorous study before deployment. Without proper understanding, we risk repeating the mistakes of AI-based decision system automation that were later found to be harmful and discriminatory (Barocas & Selbst, 2016; Miller, 2015; Angwin et al., 2016; Pagan et al., 2023; Baumann et al., 2024). Tool evaluations, simulations, and empirical studies conducted before deployment can reduce the detrimental effects of automation. Otherwise, well-known issues of LLM hallucination and bias could compromise the fairness of AIenhanced peer review (Schintler et al., 2023; Liu & Shah, 2023; Akella et al., 2025; Bonifazi et al., 2025; Zhuang et al., 2025).

Based on empirical experiments, we argue that today’s AI systems should not produce paper reviews. We ground this position in two necessary conditions that any peer review automation must satisfy:

C1. Preservation of review diversity: The system must not collapse the plurality of expert feedback that peer review aggregates. C2. Resistance to gaming: The system must not be trivially manipulable in ways that improve scores without genuine improvement of scientific content. Note: Even if these conditions were met, they would not be sufficient for full automation without deliberation on accountability, validation, and efficiency-oversight trade-offs.

We demonstrate empirically that current AI reviewers fail both conditions.

We further argue that even if those necessary conditions were fulfilled, the results would not be sufficient to automatically make fully automated AI peer review the new standard. Even a non-gameable, diversity-preserving AI system would require community deliberation on harder In this paper, we make four contributions:

  1. We demonstrate the AI reviewer hivemind effect (§3) as a failure of C1 (review diversity): AI reviewers show higher agreement within (IntraSim +8.7% to +9.8%) and across papers (InterSim +4.1% to +39.8%) than humans, both in simulation and real ICLR 2026 reviews.

  2. We introduce paper laundering (§ 4.1) as a concrete failure mode of C2 (non-gameability): zero-shot LLM rewrites boost AI review scores (+0.45, p < 0.0001) through stylistic modifications without human oversight.

  3. We show that paper laundering drives convergence toward intellectual monoculture (§ 4.2), i.e., laundered papers become significantly more similar to each other (pairwise similarity +6.5%, Cohen’s d = 1.02).

  4. We propose review diversity and non-gameability as necessary but not sufficient conditions for AI reviews (§5), and outline a science of peer review automation (§6).

Related work. Submission volumes at major AI conferences have grown significantly in recent years (Yang et al., 2026), making it increasingly difficult to find a large enough pool of qualified reviewers (Aczel et al., 2021; Shah, 2022). This imbalance forces reviewers to evaluate more papers in less time, which results in declining review quality and increased author dissatisfaction (Shah, 2022; Kuznetsov et al., 2024). The NeurIPS 2021 consistency experiment revealed a large amount of noise in human reviews, demonstrating that peer review outcomes depend a lot on reviewer assignment (Beygelzimer et al., 2021). Meanwhile, LLMassisted or fully LLM-generated reviews are already present at scale (Russo et al., 2025; Emi, 2025). These challenges have created urgent demand for solutions, making the automation of peer review processes an increasingly attractive prospect (Biswas et al., 2023; Kuznetsov et al., 2024).

AAAI 2026 provided fully LLM-generated reviews alongside human reviews. Consistent with our position, they found that participants rated AI reviews favorably on technical dimensions yet viewed them as “complementary rather than interchangeable” with human review (Biswas et al., 2026). In a large-scale survey, the AI-generated reviews were preferred on six of nine quality criteria (e.g., identifying technical errors and raising previously unconsidered points) but were also judged more likely to overemphasize minor issues and to contain technical errors of their own.

Recently, ICML 2026 introduced a two-policy framework where authors choose whether their reviewers may use LLMs for paper understanding and polishing, or not at all.2 Researchers and practitioners have explored various ways to use AI to review papers (Yuan et al., 2022; Checco et al., 2021), with recent work showing promising performance (Liang et al., 2024b; Idahl & Ahmadi, 2025). However, evaluations consistently find that LLMs correlate weakly with human judgments (Zhu et al., 2025; Shcherbiak et al., 2024), exhibit systematic score inflation (Akella et al., 2025; Li et al., 2025b; Bianchi et al., 2025b; Abdulhai et al., 2026), and fail to distinguish strong from weak papers (Bonifazi et al., 2025). Routine LLM configuration choices can themselves fabricate or suppress statistical effects in evaluation pipelines (Baumann et al., 2025), so reported AI reviewer performance can shift with undocumented setup details. Li et al. (2025a) further identified recurring weaknesses in LLM reviews, including misclassification of methodological flaws and misinterpretation of critiques. In short, while LLMs can assist human scientists, fully automating peer review raises significant fairness concerns.

Method. ICLR review data. We use all 75,800 reviews from the 19,490 papers under review at ICLR 2026. We use labels from Emi (2025), who found that 15,899 reviews (21%) are AI-generated4. We validate these labels in Appendix G.3.

AI agent reviewer simulation data. Additionally, we randomly select 60 ICLR 2026 papers, spanning a wide range of research areas. For each paper, we produce new AI reviews using the AI reviewer agents developed by (Bianchi et al., 2025b) (see Appendix B.1 for the implementation details and Appendix H for an example output). An AI review agent directly takes the paper in PDF format as input and produces a review consisting of a summary, strengths, weaknesses, questions, and a rating.

We measure C1 (diversity) with manual output inspections and the following complementary metrics.

The intra-paper inter-reviewer similarity (IntraSim) measures how similar different reviews of the same paper are. For a paper p with a set of review vector representations R(p) = {r1, . . . , rmp}, we define:

The inter-paper intra-reviewer similarity (InterSim) measures how similar reviews are across different papers. For two papers p ̸= q with review vector representation sets R(p) and R(q), we define:

Interpreting similarity. Our similarity metrics use text embeddings, which capture semantic and linguistic patterns. For both metrics, we compute cosine similarity sim between vector representations of reviews. Review embeddings are generated using OpenAI’s text-embedding-3-small model. High similarity means reviews discuss similar aspects using similar language. The value of multiple reviewers lies in noticing different things. Unlike review ratings, if two textual reviews are nearly identical, the second adds little information.

In this section, we demonstrate that AI reviewers also fail C2. They can be gamed to improve scores through fully automated paper rewriting (i.e., without any human oversight). We call this paper laundering: cosmetic paper rewrites to increase AI review scores without improving the scientific substance. We implement this process by providing the full LaTeX file together with the original AI review to an LLM in a zero-shot prompt. We compile the rewritten LaTeX code into a PDF before passing it back to the AI reviewer agents. The detailed implementation is described in Appendix B.2. Laundering one paper costs about $0.25.

Figures 4 and 5 show that paper laundering effectively games AI reviewer agents to increase paper scores. Using 60 randomly sampled ICLR 2026 papers, we apply zero-shot rewrites and compare AI review scores before and after laundering. We test 4 zero-shot prompts (including one that instructs the launderer to jailbreak the AI reviewer), 2 launderer models (GPT-5.1, GPT-5.4), and 3 reviewer models (GPT-5.1, GPT-5.4, Claude Sonnet 4.5), yielding 24 conditions.

Discussion. Human reviewers exhibit well-documented biases (Helmer et al., 2017), and the NeurIPS consistency experiments showed that roughly half of accepted papers would have received different decisions under different reviewer assignments (Beygelzimer et al., 2021). A retrospective found no correlation between reviewer scores and citation impact for accepted papers, suggesting that disagreement partly reflects genuine uncertainty papers’ future impact (Cortes & Lawrence, 2021). Collusion rings can also game the system (Littman, 2021). If human review is so flawed, why hold AI to a higher standard?

The key distinction is between distributed and centralized error. Human biases and inconsistencies are spread across multiple reviewers with different areas of expertise. Through aggregation, these errors partially cancel out. AI errors are correlated, as models trained on similar data are likely to share biases. This is an example of algorithmic monoculture (Kleinberg & Raghavan, 2021). When many decisionmakers rely on the same model, aggregate decision quality can decrease even if each individual decision looks reasonable. The same logic applies to gameability. Gaming one human reviewer does not transfer to others, so there is no universal attack. AI gameability, on the other hand, is centralized. A single rewrite strategy can boost scores across models, as we demonstrate.

Our goal is not to defend the status quo. Instead, the goal is to ensure AI-augmented peer review meets high standards, so we can build trust in peer review systems of the future.

We agree that AI can help authors improve their manuscripts in terms of readability, grammatical correctness, and other aspects. However, paper laundering, as we define it here, is a fully automated revision—without human oversight. Crucially, the changes are purely textual, without any additional experiments, and are optimized for the AI reviewer’s preferences, not for genuine substance. More importantly, AI conference guidelines require authors to take full responsibility for all paper contents, including any content generated by AI. As such, paper laundering puts them at risk of inadvertent plagiarism and scientific misconduct. Furthermore, even if textual changes improve paper clarity, the systemic effect of homogenization remains. A centralized AI-automated reviewing system would not only shape which papers are accepted but also how those papers are written.

Note that we frame non-gameability and review diversity as necessary but not sufficient conditions. If future AI systems satisfied those conditions, it would be progress. But it would not be sufficient to justify automation. Important questions of accountability and democratic legitimacy remain, which need rigorous scientific work to be addressed. Peer review plays an important role in how scientific communities collectively shape research directions. Fully delegating this functionality to AI systems would be a risky transfer of powers. Further, capable AI models may not be equally accessible to everyone. Thus, such a centralization of power cannot be based solely on conference organizer vibes.

Lastly, the likelihood that AI will get better is not a justification for deploying systems that fail today. We promote cautious automation of peer review until the risks are sufficiently well studied and addressed.

Conclusion. We demonstrate two critical failures of current AI reviewing systems and argue that they are not fit for automating peer review. First, we provide evidence that AI reviewers exhibit a hivemind effect: their outputs are far more similar than those of human reviewers, both within papers and across papers. This undermines the diversity that peer review is designed to aggregate. This comes at a measurable cost, since ratings of AI-generated reviews are less informative about final acceptance decisions than ratings of human-written reviews. Second, we show that AI review scores are trivially gameable through what we call paper laundering. This describes zero-shot paper rewrites that significantly boost scores while driving papers toward homogeneity. Our study has various limitations, including the use of a single prompt for AI reviewers and the use of third-party labels for detecting AI reviews in the wild. We discuss these in detail in Appendix A.

In this paper, we establish that current AI systems fail the necessary conditions for peer review automation. However, meeting these conditions would not automatically justify full automation. Questions of accountability, democratic legitimacy, and measurement validity require explicit community deliberation. We call for a rigorous science of peer review automation to address such questions. We envision transparent science that empirically evaluates tools before deployment, studies how humans interact with AI assistance, and develops incentive structures that maximize the value of human expertise.

The peer review crisis is real. An effective solution requires validated tools, not a simple replacement of human judgment with systems that fail to meet basic requirements.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? What human oversight must AI research systems have? How can we detect and account for LLM involvement in academic writing? Do restrictions on reviewer LLM use actually shape peer review behavior? Does AI-assisted research sacrifice exploration breadth for productivity gains?