Researchers say they don't trust AI to judge papers — but most reviewers are already quietly using it anyway.
Why do researchers resist using AI for peer review specifically?
This explores why the research community pushes back on letting AI review papers, and whether that pushback matches how reviewers actually behave.
This explores why researchers pull back from letting AI judge scientific papers, and whether that resistance matches what reviewers actually do. The first surprise is that the resistance is mostly in official policy, not in practice. A Frontiers survey of 1,645 researchers found that 53% of reviewers already use AI tools, rising to 87% among early-career researchers, mostly for drafting reports and summarizing papers How widely do peer reviewers actually use AI tools?. When ICML 2026 ran a randomized experiment that banned LLM use for some reviewers and allowed limited use for others, scores and decisions barely changed. Large numbers of reviewers also broke whichever rule they had been given Does banning LLM use in peer review change review outcomes?. So the real question is less "why won't researchers use AI?" and more "why won't they trust AI to make the decision?"
The strongest answer in the corpus is that AI reviewers fail two basic tests that any automated reviewer would need to pass. First, they think alike. Different AI systems agree with each other more than human reviewers do, and peer review depends on independent opinions. Second, they are easy to game. Rewriting a paper's wording, with no change to its science, raised AI scores by almost half a point Can AI systems safely replace human peer reviewers?. This gaming is already happening: eighteen arXiv manuscripts were found with hidden instructions telling AI reviewers to praise them, which researchers classify as research misconduct Are hidden AI prompts in preprints a deceptive research practice?. If an AI reviewer can be steered by invisible text, the gatekeeping that review exists for stops working.
Researchers also worry about what happens around review, not just inside it. A survey of 230 publications describes an arms race. AI makes it easier to produce papers, so venues automate evaluation, authors manipulate the automated evaluators, venues add defenses, and authors find ways around them Does AI create a coupled arms race in research production and review?. AI-written papers can already get past human review at the workshop level: one of three fully AI-generated papers from Sakana's AI Scientist scored above the acceptance line at an ICLR 2025 workshop. Its authors later found a citation error and decided none of the three met main-conference standards Can AI-generated papers pass peer review undetected? Can AI systems generate research papers that pass peer review?. If machines both write and judge the papers, reviewers fear that no human ever checks whether the work is actually right.
The pressure runs the other way too, and this is where the debate gets interesting. One argument holds that if we accept AI-accelerated research, we have already committed to AI-assisted review, because human reviewers cannot keep up. The question becomes how much to hand over while keeping humans accountable Can human review keep pace with AI-accelerated research generation?. Nature makes a similar case that funders and publishers need policies now, before review workloads overwhelm the system Can AI-generated research outpace peer review systems?. Some groups are already building separate venues where AI-generated work goes through repeated automated review-and-revise rounds with built-in defenses against hidden prompts Can automated review loops handle AI-generated research at scale?.
What you might not expect is that some of the pushback has little to do with AI. Position papers argue that peer review at AI conferences was already struggling. Authors, reviewers, and venues all share the blame, with measurable biases such as scores tracking review length Can two-stage review and badges fix AI conference peer review?. Munger goes further and argues that the bottleneck is the peer-reviewed PDF itself and the academic incentives built around it, not what AI can or can't do Can AI help social science move beyond the peer-reviewed PDF?. Read this way, resisting AI review is partly a way of defending a fragile system that AI puts under more strain.
Sources 12 notes
Frontiers' May-June 2025 survey of 1,645 researchers found 53% of reviewers use AI tools, with adoption reaching 87% among early-career researchers. Most use AI for drafting reports or summarizing findings, and researchers express desire for clearer policies to guide more advanced applications.
A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Eighteen arXiv manuscripts contained concealed instructions directing AI reviewers to give positive assessments. The practice qualifies as questionable research conduct because concealment plus self-serving design violates ethics regardless of stated intent.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
Show all 12 sources
Sakana AI's end-to-end system produced a paper that scored 6.33 in double-blind ICLR 2025 workshop review, meeting acceptance thresholds, but was withdrawn under pre-agreed protocol. Authors later identified a citation error and judged none of three submissions suitable for main-track publication.
AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.
The PAT framework argues that accepting AI-driven research output commits us structurally to AI-assisted verification, not as option but necessity. A taxonomy of four collaboration levels—from author tools to reviewer augmentation—provides the governance scaffolding to manage this transition while keeping humans accountable.
A Nature editorial argues AI science has moved from preprint novelty to published output, requiring immediate institutional, funder, and publisher policies on authorship, credit, and review workload before systems are overwhelmed.
aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Munger contends that the peer-reviewed PDF combines distinct functions—archive, literature review, theory, methods, results—that AI could separate into recombinable forms. He identifies document form and academic incentives, not AI capability limits, as the bottleneck preventing this transition.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- The AI Scientist Generates its First Peer-Reviewed Scientific Publication
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI for Auto-Research: Roadmap & User Guide