Could AI catch the mistakes peer reviewers miss, if it's given time to check each proof and experiment step by step?
Could AI improve peer review rigor and catch human-missed errors?
This explores whether AI can make peer review more rigorous, especially by catching mistakes in proofs, experiments, and arguments that human reviewers let slip through, and where that help stops.
This explores whether AI can make peer review more rigorous by catching mistakes human reviewers miss, and where that help runs out. The corpus gives two answers. As a careful checker working alongside humans, AI already shows real promise. As a stand-alone judge of what deserves publication, it fails in ways that are easy to miss.
The strongest evidence for AI catching errors comes from giving the model more time to think. An agentic reviewer that spends extra compute at test time checking proofs and experiments line by line found 34% more math errors than a model reviewing in one quick pass. It also flagged serious flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. The lesson is that the speed of a single response isn't the limit. Slow, step-by-step checking is. A related idea from automated paper writing points the same way: keep the model's judgment separate from checks a program can run with a guaranteed result, so that reliability doesn't depend on the model being right Can separating judgment from verification improve research paper reliability?. Closed loops of automated review followed by revision have also measurably improved AI-generated papers Can automated review loops handle AI-generated research at scale?.
AI can also make the human reviewers better, which may be the more surprising result. In a randomized trial at ICLR 2025, optional feedback from Claude-based agents led 27% of reviewers to revise their reviews. Blinded raters judged the revised reviews more informative and clearer Can LLM feedback help peer reviewers improve their own reviews?. In this setup the AI reviews the review, not the paper. That fits a broader argument that review failures are shared among authors, reviewers, and venues, and that fixes should target known biases, such as scores tracking how long a review is Can two-stage review and badges fix AI conference peer review?.
The warning signs show up when AI is asked to be the judge rather than the checker. AI reviewers agree with each other far more than human reviewers do, a 'hivemind' effect that removes the variety of viewpoints peer review relies on. They are also easy to game: rewording a paper with no change to its science raised AI scores by 0.45 points Can AI systems safely replace human peer reviewers?. Fully AI-generated papers have already cleared workshop review. One scored 6.33 at an ICLR 2025 workshop, yet its own authors later found a citation error and judged none of their three submissions ready for the main conference Can AI-generated papers pass peer review undetected? Can AI systems generate research papers that pass peer review?. So the reviewers being fooled were human, and that is exactly why some argue AI help is now required. If AI speeds up how fast papers get written, human-only review can't keep up Can human review keep pace with AI-accelerated research generation?.
The part you may not have expected: this is less a tool question than an arms race. A survey of 230 publications describes six linked dynamics: more papers being produced, more automated evaluation, manipulation, defenses, evasion of those defenses, and feedback across the whole system. Each move provokes a response, and the evidence gets thinner the further out you look Does AI create a coupled arms race in research production and review?. One more twist: models fine-tuned on where papers were actually published predicted outcomes better than expert reviewers did Can institutional publication records train better scientific evaluators?. But predicting where a paper will land is not the same as checking whether it is correct. A model trained on past decisions may simply learn the field's habits, prestige signals included. The practical takeaway from the corpus: let AI do the slow verification humans skip, and keep humans accountable for the judgment.
Sources 11 notes
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.
aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Show all 11 sources
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.
The PAT framework argues that accepting AI-driven research output commits us structurally to AI-assisted verification, not as option but necessity. A taxonomy of four collaboration levels—from author tools to reviewer augmentation—provides the governance scaffolding to manage this transition while keeping humans accountable.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
LLMs fine-tuned on eight social science publication records beat both expert majority votes and frontier reasoning models at evaluating research pitches, reaching 59.2% accuracy in management versus 41.6% expert agreement. The models learned field-level evaluation logic from institutional stratification rather than written criteria.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI for Auto-Research: Roadmap & User Guide
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- The AI Scientist Generates its First Peer-Reviewed Scientific Publication