Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026

Paper · arXiv 2609.19420 · Published September 16, 2026
LLM Evaluations and Benchmarks

LLMs are rapidly reshaping peer review, making it important to understand how reviewers use them in practice and how different LLM-use policies affect review outcomes. We investigate these questions through a randomized experiment and an anonymous post-survey at ICML 2026, a major machine learning conference involving over 24,000 papers and 17,000 reviewers. Reviewers were assigned to either a conservative policy prohibiting all LLM use or a permissive policy allowing limited assistance, with randomization among a subset of main-track papers and reviewers. Policy assignment had near-zero effects on final paper decisions, paper scores, and reviewer confidence, although reviews under the permissive policy were 5.5-7% longer. Post-survey responses (N=1,486) revealed diverse attitudes toward LLMs and substantial noncompliance: 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. We discuss implications for future peer-review policy and tool design.

Introduction. Rapidly growing submission volumes are increasing the burden on peer-review systems and amplifying longstanding concerns about “reviewer fatigue” [67]. At CHI and ICML, two leading conferences in human-computer interaction (HCI) and machine learning (ML), the increase in submissions between 2023 and 2026 was 3,182 to 6,730 (+112%) and 6,538 to 24,661 (+277%), respectively. In response to these pressures and the growing capabilities of large language models (LLMs), it has become common for researchers to incorporate the use of LLMs into their peer review practices [50, 52–54]. Computer Science (CS) conferences, in turn, have begun to publish explicit peer-review LLM policies and explore the use of LLM assistance in the review workflow [1, 2, 7, 22, 33–36, 39, 59, 83].1 (See Section A for an overview of these policies.)

Discussion / Conclusion. 6.1 Implications for the Design of Policies and Tools for Peer Review In this section, we reflect on our findings and discuss promising directions for future peer-review policies and tools. 6.1.1 Design Realistic Policies and Cultivate the Right Environment. Traditionally, reviewing has been a valued, voluntary activity. In a survey of 307 CHI reviewers, Nobarany et al. [61] found that “encouraging high-quality research, giving back to the research community, and finding out about new research” were reviewers’ primary motivations for reviewing. Our post-survey responses echoed these themes, describing science as a “community process” and emphasizing that “the fundamental purpose of peer review [...] is to provide expert, accountable, and independent judgment.” These findings suggest that, with the right culture and environment in place, reviewers would be motivated to conduct high-quality reviews. However, our findings suggest that these conditions are not in place, as evidenced by the high rates of noncompliance in both the Pangram analysis (Section 4.2.3) and post-survey results (Section 5.2.1).

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why does verification consistently lag behind AI generation? Does AI text rewriting systematically distort writer intent and preference? How should human oversight be integrated with autonomous AI systems? When should tasks involve human-AI partnership versus full automation? Why do readers trust citations and complexity regardless of accuracy? Can AI-generated outputs constitute genuine knowledge or valid claims? How do evaluation biases undermine LLM quality assessment systems? Why do LLM research ideas score high on novelty yet collapse into low diversity? Can ensemble evaluation methods reduce bias more than single judges? Why can LLMs generate ideas better than they evaluate them? Can model confidence signals reliably improve reasoning quality and calibration? How do language models inherit human biases from training data?