SYNTHESIS NOTE
Topics›Evaluations›this note

Does banning LLM use in peer review change review outcomes?

Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.

Synthesis note · 2026-09-25 · sourced from Evaluations

The paper reports a randomized experiment at ICML 2026, a conference with over 24,000 papers and 17,000 reviewers. Among a subset of main-track papers and reviewers, some reviewers were assigned a "conservative policy prohibiting all LLM use" and others a "permissive policy allowing limited assistance." Policy assignment had "near-zero effects on final paper decisions, paper scores, and reviewer confidence." The one measured difference was length: reviews under the permissive policy were 5.5-7% longer.

The second finding is how far reviewers stepped outside the policy they received. In an anonymous post-survey (N=1,486), 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. The discussion adds that a Pangram analysis pointed the same way, and it treats the "high rates of noncompliance" as evidence that the conditions for high-quality reviewing are not in place.

The paper's reasoning starts from load. ICML submissions grew from 6,538 to 24,661 (+277%) between 2023 and 2026, and CHI's grew 112%, which amplifies "reviewer fatigue." The discussion then contrasts this with the traditional picture of reviewing as a valued, voluntary activity, motivated by giving back to the community. Survey respondents describe science as a "community process" and peer review as "expert, accountable, and independent judgment." From that the authors argue for realistic policies and a supportive environment. A rule written for the reviewer culture the field remembers, rather than the workload it now has, is being ignored by a sizable minority.

This is field evidence for a transition that Can human review keep pace with AI-accelerated research generation? frames normatively. Tool-for-reviewers use is already happening at scale, sanctioned or not, and neither a ban nor limited permission visibly changed the outcomes measured here. It also contrasts with Can structured pipelines make LLM novelty assessment reliable? and Can inference scaling help reviewers catch errors humans miss?. Those notes concern designed pipelines evaluated against human judgment. This paper concerns unstructured, self-directed use by reviewers, and its outcome measures are scores, decisions and confidence rather than agreement or flaw detection.

The excerpt is silent on several things that limit the reading. It gives no confidence intervals, no size for the randomized subset, and no measure of review quality or accuracy, so "near-zero" effects on scores do not show that reviews were equally good. The noncompliance figures are self-reports from 1,486 respondents out of roughly 17,000 reviewers, and the excerpt does not give the response rate or say what the Pangram analysis found. It also does not say how noncompliance affected the contrast between the two arms. What the evidence supports is narrower: as a lever on the measured outcomes, the choice between a ban and limited permission looks weak, and any policy should assume partial noncompliance.

Inquiring lines that read this note 4

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What safeguards enable trustworthy AI-assisted scientific peer review at scale? Why do LLM recommenders underperform collaborative filtering despite their capabilities?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 104 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

banning versus limiting LLM use in peer review had near-zero effects on scores and decisions — noncompliance under both policies was substantial