INQUIRING LINE

AI can now write papers faster than people can review them, so who checks the flood of research?

How much has peer review workload grown at major conferences?

This explores how much the reviewing burden (submission counts, papers per reviewer) has grown at top AI conferences, and the corpus has no submission or reviewer-load figures, so it can only speak to the pressure around that question.


This explores how much peer review workload has grown at major conferences. The corpus has no numbers on it: no submission counts by year, no papers-per-reviewer figures, no hours spent. Any figure I gave would be invented. What it does hold is material on why the load is expected to keep growing and on what conferences are trying in response.

The clearest statement of the pressure is a framework arguing that if AI speeds up paper generation, review has to be automated too, or the pipeline collapses. The argument is structural: accepting AI-driven research output commits a field to AI-assisted verification. The authors offer a four-level taxonomy of human-AI collaboration, from author-side tools to reviewer augmentation, so humans stay accountable while the load shifts Can human review keep pace with AI-accelerated research generation?. That is a claim about where the load is heading, not a measurement of where it is. A related proposal says AI-generated research lacks a suitable venue at all, and describes a venue built around automated review-and-refine loops Can automated review loops handle AI-generated research at scale?.

The corpus also shows what reviewers do when the load is high. An ICML 2026 randomized experiment compared banning LLM use in peer review with allowing limited use. Paper scores, decisions and reviewer confidence barely moved, and substantial fractions of reviewers broke whichever rule they were assigned Does banning LLM use in peer review change review outcomes?. The note doesn't say why they broke the rules, so I wouldn't read it as proof of overload. It does suggest that policy alone isn't controlling how reviews get produced.

On the response side, the notes describe tools aimed at the review bottleneck. A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment with human reviewers on 182 ICLR submissions when judging novelty Can structured pipelines make LLM novelty assessment reliable?. An agentic reviewer that spends extra compute checking proofs line by line surfaced flaws at STOC and ICML that had passed human review Can inference scaling help reviewers catch errors humans miss?. So the collection frames the load as a scaling problem and shows early attempts to automate parts of it, but the size of the growth is a gap. To get real figures you would need conference submission and reviewer-assignment statistics from outside this library.


Sources 5 notes

Can human review keep pace with AI-accelerated research generation?

The PAT framework argues that accepting AI-driven research output commits us structurally to AI-assisted verification, not as option but necessity. A taxonomy of four collaboration levels—from author tools to reviewer augmentation—provides the governance scaffolding to manage this transition while keeping humans accountable.

Can automated review loops handle AI-generated research at scale?

aiXiv demonstrates that iterative review-refine cycles with automated retrieval-augmented evaluation and prompt-injection defenses measurably enhance proposal and paper quality, addressing the structural gap where AI-generated research lacks appropriate publication venues.

Does banning LLM use in peer review change review outcomes?

A randomized experiment at ICML 2026 found that prohibiting LLM use versus allowing limited use barely changed paper scores, decisions, or reviewer confidence. Meanwhile, substantial fractions of reviewers broke whichever rule they were given.

Can structured pipelines make LLM novelty assessment reliable?

A three-stage pipeline (extract claims, retrieve related work, compare) reached 86.5% reasoning alignment and 75.3% conclusion agreement with human reviewers on 182 ICLR submissions, outperforming holistic LLM baselines.

Can inference scaling help reviewers catch errors humans miss?

PAT, an agentic reviewer using test-time compute to check proofs and experiments line by line, achieves 34% better recall on math errors than zero-shot approaches and surfaced critical flaws at STOC and ICML that passed human review.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.