INQUIRING LINE

One AI review can look as good as a human's, but could AI replace every peer reviewer, and what would break?

Could AI feedback work as a substitute for human peer review entirely?

This explores whether AI-generated reviews could fully replace human peer reviewers, rather than assist them, and what the corpus shows about where AI review holds up and where it breaks.


This explores whether AI could take over peer review completely, not just help with it. The corpus suggests that a single AI review can look as good as a single human one, but swapping out every human reviewer breaks something that no individual review measures. The best evidence for replacement is a study of thousands of Nature and ICLR papers. It found that GPT-4's comments overlapped with a given human reviewer's about as often as two human reviewers overlapped with each other, and most of the researchers surveyed found the feedback useful Can GPT-4 feedback match what human reviewers catch?. On some tasks AI does better than people. An agentic reviewer that spends extra compute checking proofs and experiments line by line found serious flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?.

The catch is that peer review is valuable as a system, not just review by review. Its value depends on reviewers disagreeing independently and on being hard to manipulate, and AI reviewers fail on both counts Can AI systems safely replace human peer reviewers?. AI models agree with each other more than humans do, a 'hivemind' effect, so three AI reviewers behave more like one reviewer counted three times. They're also easy to game: simply rewriting a paper's text raised AI scores by about half a point without changing the science at all. If every venue used AI reviewers, authors would learn to write for the model, and the review signal would turn into a measure of how well a paper is optimized for that model.

The 'AI papers passed peer review' headlines are worth reading carefully, because they test a different question. Fully AI-generated manuscripts have cleared workshop review, but only one of three did. Its own authors withdrew it, later found a citation error, and judged that none of the three met main-conference standards Can AI systems generate research papers that pass peer review? Can AI-generated papers pass peer review undetected?. The original AI Scientist also graded its own work with AI reviewers Can one AI system complete a full research cycle end-to-end?. That closed loop is exactly where the hivemind problem would compound. Projects that build venues around automated review-and-revise cycles, like aiXiv, treat defenses against prompt injection as a core requirement Can automated review loops handle AI-generated research at scale?. This fits a wider survey describing AI paper production and AI review as a coupled arms race of manipulation, defenses, and evasion Does AI create a coupled arms race in research production and review?.

The strongest results so far come from AI reviewing the reviewers. In a randomized trial at ICLR 2025, AI feedback on draft reviews led 27% of reviewers to revise them, and blinded raters judged the revised reviews more informative Can LLM feedback help peer reviewers improve their own reviews?. A separate ICML experiment found that banning LLM use versus allowing limited use barely changed scores or decisions, partly because many reviewers ignored whichever rule they were given Does banning LLM use in peer review change review outcomes?. That suggests the live question is how AI and humans divide the work, since banning AI barely changes outcomes.

There's one serious argument that some automation is unavoidable. If AI speeds up paper production, human reviewers can't keep up, so some level of AI-assisted checking becomes necessary. Even so, its proponents describe graded levels of AI collaboration that keep humans accountable Can human review keep pace with AI-accelerated research generation?. Others argue that peer review's problems come from authors, reviewers, and venues alike, and need fixes like letting authors rate review quality, which no AI reviewer provides Can two-stage review and badges fix AI conference peer review?. The overall picture: AI is already a capable reviewer, and a strong fact-checker for proofs and experiments. It can't yet supply the independent disagreement that makes peer review work.


Sources 12 notes

Can GPT-4 feedback match what human reviewers catch?

A study of 3,096 Nature papers and 1,709 ICLR papers found GPT-4 matched individual reviewers' points 30.85% of the time versus 28.58% for two human reviewers. Fifty-seven percent of surveyed researchers found the feedback helpful.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Can AI systems safely replace human peer reviewers?

AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.

Can AI systems generate research papers that pass peer review?

AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.

Can AI-generated papers pass peer review undetected?

Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.

Show all 12 sources
Can one AI system complete a full research cycle end-to-end?

The AI Scientist performed ideation, coding, experiments, writing, and self-review autonomously, producing a manuscript that passed the first round at a machine learning workshop with 70% acceptance rate. Five ensemble reviewers and an area-chair model judged the output against NeurIPS guidelines.

Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Does AI create a coupled arms race in research production and review?

A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.

Can LLM feedback help peer reviewers improve their own reviews?

A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Can human review keep pace with AI-accelerated research generation?

The PAT framework argues that accepting AI-driven research output commits us structurally to AI-assisted verification, not as option but necessity. A taxonomy of four collaboration levels—from author tools to reviewer augmentation—provides the governance scaffolding to manage this transition while keeping humans accountable.

Can two-stage review and badges fix AI conference peer review?

Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.