AI clearly changes what peer reviewers write, but in controlled trials it barely moved scores or which papers got accepted.
Did adding AI reviews actually change peer review decisions or paper outcomes?
This asks whether bringing AI into peer review, as feedback to reviewers, as a tool reviewers use, or as a reviewer in its own right, has measurably changed what gets accepted, how papers score, or what happens to them afterward, rather than just changing how reviews read.
This asks whether AI in peer review has changed actual outcomes, meaning scores, accept/reject decisions and paper fates, rather than just the wording of reviews. The short answer from the collection is surprising: AI clearly changes the reviews themselves, but the controlled evidence so far shows almost no effect on decisions. The best test comes from ICML 2026, which randomly assigned reviewers either a ban on LLM use or permission for limited use. Paper scores, decisions and reviewer confidence barely moved between the two groups. A large share of reviewers also broke whichever rule they were given, which muddies what the 'ban' group really measured Does banning LLM use in peer review change review outcomes?.
Where AI did visibly change something, the change was in the text of reviews. In a randomized trial at ICLR 2025, reviewers could get optional feedback on their drafts from a Claude-based agent, and 27 percent of them revised their reviews. Blinded raters judged the revised reviews more specific and informative Can LLM feedback help peer reviewers improve their own reviews?. The summary here reports better reviews, not different verdicts. A clearer critique and a different acceptance decision are separate things, and the collection's evidence covers the first more than the second. This is happening at scale: a Frontiers survey of 1,645 researchers found 53 percent of reviewers already use AI tools, and 87 percent of early-career researchers do, mostly for drafting and summarizing How widely do peer reviewers actually use AI tools?.
The collection also suggests where AI could change outcomes, for better or worse. On the useful side, PAT, an agentic reviewer that spends extra computation checking proofs and experiments line by line, found serious flaws in papers that had already passed human review at STOC and ICML Can inference scaling help reviewers catch errors humans miss?. That kind of catch could flip a decision if it arrived in time. On the risky side, AI reviewers agree with each other more than humans do. Rewriting a paper's prose, with no change to its science, raised AI scores by 0.45 points Can AI systems safely replace human peer reviewers?. Authors have noticed: 18 arXiv manuscripts were found with hidden instructions telling AI reviewers to give positive assessments Are hidden AI prompts in preprints a deceptive research practice?. If AI reviews carried real weight in decisions, these weaknesses would become ways to game acceptance.
The outcome question also runs the other way: can AI-written papers get through review? Sakana's AI Scientist-v2 submitted three fully automated papers to an ICLR 2025 workshop. One averaged 6.33 and met the acceptance threshold, but it was withdrawn under a protocol agreed in advance. The authors later found a citation error and judged that none of the three papers met main-conference standards Can AI systems generate research papers that pass peer review? Can AI-generated papers pass peer review undetected?.
The collection says its own answer is incomplete. A survey of 230 publications describes AI in research and review as an arms race of production, automated evaluation, manipulation and defense. It finds the strongest evidence for early effects, such as adoption and changes in review text, and the weakest for long-run effects on outcomes and the wider research system Does AI create a coupled arms race in research production and review?. So the honest answer is that AI has changed the reviewing process, but the randomized evidence here does not show it changing which papers get in. Proposals to let authors rate reviews before seeing verdicts suggest the field may care more about review quality than verdicts Can two-stage review and badges fix AI conference peer review?.
Sources 10 notes
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
Frontiers' May-June 2025 survey of 1,645 researchers found 53% of reviewers use AI tools, with adoption reaching 87% among early-career researchers. Most use AI for drafting reports or summarizing findings, and researchers express desire for clearer policies to guide more advanced applications.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Show all 10 sources
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv