Peer review's workload problem may feed itself: overloaded reviewers lower accuracy, which tempts authors to send more speculative papers.
How fast is scientific publishing growing relative to reviewer capacity?
This explores whether the collection has hard numbers on how quickly paper submissions are outpacing the supply of qualified peer reviewers, and what happens to review when they do.
This explores how fast paper submissions are growing compared with the number of people available to review them. The short answer: the collection doesn't have that ratio. None of these notes gives growth rates for submissions or for the reviewer pool. What the corpus does have is more useful in a way. It explains why the gap tends to widen once it opens, and it shows what reviewers are already doing to cope.
The central idea is a feedback loop. A two-journal model shows that rising submissions overload unpaid reviewers. Journals then either recruit less qualified reviewers or pile more work on the existing ones. Review accuracy falls, and when review gets noisier, authors have more reason to submit speculative work and hope it gets through. That pushes submissions up again Does peer review quality collapse under submission overload?. So the problem isn't one curve pulling ahead of another. Overload makes overload worse. The authors are candid that they haven't measured how strong this effect is in real data. A survey of 230 publications sees a similar loop across the field: AI scales up paper production, review gets automated in response, then come manipulation, defenses and evasion. The evidence is strongest for the early stages and thins out for the long-term feedback Does AI create a coupled arms race in research production and review?.
The best signs of strain come from what reviewers are doing. A Frontiers survey of 1,645 researchers found that 53% of reviewers use AI tools, and 87% of early-career reviewers do, mostly to draft reports or summarize papers How widely do peer reviewers actually use AI tools?. At ICLR, Pangram Labs estimated that 21% of reviews were written entirely by AI. Reviews with more AI text gave systematically higher scores How much AI content appears in peer review at ICLR?. This is what capacity shortfall looks like on the ground: reviewers handing off part of the job, and the handoff quietly changing the verdicts. Meanwhile, generation is speeding up. One fully AI-written paper scored 6.33 at an ICLR workshop, high enough to be accepted, though its authors said it wasn't ready for the main conference Can AI systems generate research papers that pass peer review?.
From there the debate splits. One side argues that if AI speeds up writing papers, it has to speed up reviewing them too, or the pipeline collapses. Under that view, the open question is how to keep humans accountable Can human review keep pace with AI-accelerated research generation?. Supporting evidence: an agentic reviewer that checks proofs line by line found errors in STOC and ICML papers that human reviewers missed Can inference scaling help reviewers catch errors humans miss?. A venue called aiXiv shows that repeated automated review-and-revise rounds improve AI-generated papers Can automated review loops handle AI-generated research at scale?. The opposing finding is sobering. AI reviewers agree with each other far more than human reviewers do, which the authors call a "hivemind." And a simple rewrite of a paper's text raised AI scores by 0.45 points without improving the science Can AI systems safely replace human peer reviewers?. Automating review at scale could swap a shortage of reviewers for a monoculture of reviewers that is easy to game.
The corpus offers one more angle. Capacity isn't only about how many reviewers there are. It's also about whether good reviewing gets rewarded. One position paper argues that authors, reviewers and venues all share the blame. It proposes letting authors rate the quality of a review before they see the decision, plus badges that reward thorough reviewers Can two-stage review and badges fix AI conference peer review?. If you came for a growth chart, the collection points elsewhere. The worry is less that reviewers are falling behind and more that the ways they keep up change what peer review measures.
Sources 10 notes
A two-journal model shows that rising submissions overtax unpaid reviewers, forcing journals to recruit less qualified reviewers or overload existing ones, which drops review accuracy and incentivizes authors to submit more speculatively, driving submissions higher. The mechanism is structural but its empirical strength remains to be measured.
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
Frontiers' May-June 2025 survey of 1,645 researchers found 53% of reviewers use AI tools, with adoption reaching 87% among early-career researchers. Most use AI for drafting reports or summarizing findings, and researchers express desire for clearer policies to guide more advanced applications.
Pangram Labs' analysis of ICLR's public review corpus estimates 21% of reviews were fully AI-generated and over half had some AI involvement. Reviews with more AI text received systematically higher scores, suggesting AI may amplify positive bias rather than just rephrase human judgment.
AI Scientist-v2 submitted three fully autonomous manuscripts to ICLR; one averaged 6.33 from reviewers and ranked in the top 45% of workshop submissions. The authors acknowledged the work does not yet meet top-tier conference standards and withdrew the accepted paper before publication.
Show all 10 sources
The PAT framework argues that accepting AI-driven research output commits us structurally to AI-assisted verification, not as option but necessity. A taxonomy of four collaboration levels—from author tools to reviewer augmentation—provides the governance scaffolding to manage this transition while keeping humans accountable.
PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.
aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.
AI systems show a hivemind effect, agreeing more with each other than humans do across papers. Zero-shot rewrites of paper text raise AI scores by 0.45 points without improving scientific content, demonstrating trivial gameability at scale.
Authors, reviewers, and venues all contribute to peer review failures at major AI conferences. A proposed two-stage system lets authors rate review quality before seeing verdicts, and a badge system rewards reviewer thoroughness, targeting measured biases like rating-length correlation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- AI for Auto-Research: Roadmap & User Guide
- The AI Scientist Generates its First Peer-Reviewed Scientific Publication
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap