Told to avoid AI or limit it, many peer reviewers broke the rule, yet review scores came out nearly the same either way.
Do peer reviewers actually follow restrictions on using AI tools themselves?
This explores whether reviewers obey the AI-use rules that conferences and journals set for them, and what that compliance gap says about whether such rules are the right lever at all.
This explores whether peer reviewers stick to the AI-use rules venues give them, and what happens when they don't. The most direct evidence is blunt. In a randomized experiment at ICML 2026, some reviewers were told not to use LLMs at all and others were allowed limited use. Substantial fractions of reviewers broke whichever rule they were given Does banning LLM use in peer review change review outcomes?. The surprising part is that it barely mattered. Paper scores, accept/reject decisions and reviewer confidence came out almost the same under both policies. So the honest answer is: often not, and the rule itself may not be what shapes the outcome.
The broader picture explains why rules are hard to enforce. A Frontiers survey of more than 1,600 researchers found that over half of reviewers already use AI tools, rising to 87% among early-career researchers. Most use them for drafting reports and summarizing papers, and many said they want clearer policies, not stricter ones How widely do peer reviewers actually use AI tools?. When AI use is this widespread and this ordinary, a ban starts to look like a rule people treat as optional. Authors seem to have reached the same conclusion. Eighteen arXiv manuscripts were found with hidden instructions telling AI reviewers to rate the paper favorably Are hidden AI prompts in preprints a deceptive research practice?. Nobody plants a trap for AI reviewers unless they expect AI to be reading.
The restrictions aren't arbitrary, though. AI reviewers show a 'hivemind' effect: different models agree with each other more than human reviewers do. They are also easy to game. Simply rewriting a paper's text raised AI scores by almost half a point without improving the science Can AI systems safely replace human peer reviewers?. Hidden-prompt manipulation, defenses against it and the next round of evasion form what one survey of 230 publications calls a coupled arms race between paper production and paper review Does AI create a coupled arms race in research production and review?.
This suggests the real question isn't ban versus allow. It's what kind of AI involvement venues want to design for. ICLR 2025 tried a sanctioned approach: optional AI feedback on reviewers' drafts. That led 27% of reviewers to revise, and blinded raters judged the revised reviews more informative Can LLM feedback help peer reviewers improve their own reviews?. Others argue that if AI speeds up paper writing, review has to be AI-assisted to keep up. They propose formal levels of human-AI collaboration that keep humans accountable Can human review keep pace with AI-accelerated research generation?. A separate position paper argues review failures are shared by authors, reviewers and venues, and that reviewer incentives need redesigning, not just policing Can two-stage review and badges fix AI conference peer review?.
The takeaway: the ICML experiment hints that reviewer behavior has already moved past the policy. A rule many reviewers ignore, and that doesn't change outcomes when followed, may matter less than how AI help is built into the review process openly. The collection has one solid compliance study, so treat 'substantial fractions' as a strong signal rather than a settled rate across venues.
Sources 8 notes
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
Frontiers' May-June 2025 survey of 1,645 researchers found 53% of reviewers use AI tools, with adoption reaching 87% among early-career researchers. Most use AI for drafting reports or summarizing findings, and researchers express desire for clearer policies to guide more advanced applications.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
Show all 8 sources
A randomized trial at ICLR 2025 found that optional, gated feedback from Claude-based agents led over a quarter of reviewers to update their reviews, incorporating suggestions that blinded raters judged as more informative and clear.
The PAT framework argues that accepting AI-driven research output commits us structurally to AI-assisted verification, not as option but necessity. A taxonomy of four collaboration levels—from author tools to reviewer augmentation—provides the governance scaffolding to manage this transition while keeping humans accountable.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Position: The AI Conference Peer Review Crisis Demands Author Feedback and Reviewer Rewards
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv