Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
The paper reports a randomized experiment at ICML 2026, a conference with over 24,000 papers and 17,000 reviewers. Among a subset of main-track papers and reviewers, some reviewers were assigned a "conservative policy prohibiting all LLM use" and others a "permissive policy allowing limited assistance." Policy assignment had "near-zero effects on final paper decisions, paper scores, and reviewer confidence." The one measured difference was length: reviews under the permissive policy were 5.5-7% longer.
The second finding is how far reviewers stepped outside the policy they received. In an anonymous post-survey (N=1,486), 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. The discussion adds that a Pangram analysis pointed the same way, and it treats the "high rates of noncompliance" as evidence that the conditions for high-quality reviewing are not in place.
The paper's reasoning starts from load. ICML submissions grew from 6,538 to 24,661 (+277%) between 2023 and 2026, and CHI's grew 112%, which amplifies "reviewer fatigue." The discussion then contrasts this with the traditional picture of reviewing as a valued, voluntary activity, motivated by giving back to the community. Survey respondents describe science as a "community process" and peer review as "expert, accountable, and independent judgment." From that the authors argue for realistic policies and a supportive environment. A rule written for the reviewer culture the field remembers, rather than the workload it now has, is being ignored by a sizable minority.
This is field evidence for a transition that Can human review keep pace with AI-accelerated research generation? frames normatively. Tool-for-reviewers use is already happening at scale, sanctioned or not, and neither a ban nor limited permission visibly changed the outcomes measured here. It also contrasts with Can structured pipelines make LLM novelty assessment reliable? and Can inference scaling help reviewers catch errors humans miss?. Those notes concern designed pipelines evaluated against human judgment. This paper concerns unstructured, self-directed use by reviewers, and its outcome measures are scores, decisions and confidence rather than agreement or flaw detection.
The excerpt is silent on several things that limit the reading. It gives no confidence intervals, no size for the randomized subset, and no measure of review quality or accuracy, so "near-zero" effects on scores do not show that reviews were equally good. The noncompliance figures are self-reports from 1,486 respondents out of roughly 17,000 reviewers, and the excerpt does not give the response rate or say what the Pangram analysis found. It also does not say how noncompliance affected the contrast between the two arms. What the evidence supports is narrower: as a lever on the measured outcomes, the choice between a ban and limited permission looks weak, and any policy should assume partial noncompliance.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What safeguards enable trustworthy AI-assisted scientific peer review at scale? Why do LLM recommenders underperform collaborative filtering despite their capabilities?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can human review keep pace with AI-accelerated research generation?
As AI systems generate hypotheses, code, and proofs faster than humans can verify them, does the bottleneck at peer review force verification itself to become automated? What governance structures enable this transition safely?
normative taxonomy of collaboration levels; this paper observes reviewer-side LLM use already occurring under two different policies
-
Can structured pipelines make LLM novelty assessment reliable?
Explores whether breaking novelty assessment into extraction, retrieval, and comparison stages helps LLMs align with human peer reviewers and produce more rigorous, evidence-based evaluations.
designed pipeline measured for alignment; here reviewers' own uncontrolled use is measured for score and decision effects
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
purpose-built reviewing tool, where this paper studies informal reviewer use and the policies governing it
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
- How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
- When Reject Turns into Accept: Quantifying the Vulnerability of LLM-Based Scientific Reviewers to Indirect Prompt Injection
- AI Meets the Classroom: When Does ChatGPT Harm Learning?
- Beyond "Not Novel Enough": Enriching Scholarly Critique with LLM-Assisted Feedback
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
- Using Large Language Models to Create AI Personas for Replication and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings
- The Impossibility of Fair LLMs
Original note title
banning versus limiting LLM use in peer review had near-zero effects on scores and decisions — noncompliance under both policies was substantial