Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
LLMs are rapidly reshaping peer review, making it important to understand how reviewers use them in practice and how different LLM-use policies affect review outcomes. We investigate these questions through a randomized experiment and an anonymous post-survey at ICML 2026, a major machine learning conference involving over 24,000 papers and 17,000 reviewers. Reviewers were assigned to either a conservative policy prohibiting all LLM use or a permissive policy allowing limited assistance, with randomization among a subset of main-track papers and reviewers. Policy assignment had near-zero effects on final paper decisions, paper scores, and reviewer confidence, although reviews under the permissive policy were 5.5-7% longer. Post-survey responses (N=1,486) revealed diverse attitudes toward LLMs and substantial noncompliance: 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. We discuss implications for future peer-review policy and tool design.
Introduction. Rapidly growing submission volumes are increasing the burden on peer-review systems and amplifying longstanding concerns about “reviewer fatigue” [67]. At CHI and ICML, two leading conferences in human-computer interaction (HCI) and machine learning (ML), the increase in submissions between 2023 and 2026 was 3,182 to 6,730 (+112%) and 6,538 to 24,661 (+277%), respectively. In response to these pressures and the growing capabilities of large language models (LLMs), it has become common for researchers to incorporate the use of LLMs into their peer review practices [50, 52–54]. Computer Science (CS) conferences, in turn, have begun to publish explicit peer-review LLM policies and explore the use of LLM assistance in the review workflow [1, 2, 7, 22, 33–36, 39, 59, 83].1 (See Section A for an overview of these policies.)
Discussion / Conclusion. 6.1 Implications for the Design of Policies and Tools for Peer Review In this section, we reflect on our findings and discuss promising directions for future peer-review policies and tools. 6.1.1 Design Realistic Policies and Cultivate the Right Environment. Traditionally, reviewing has been a valued, voluntary activity. In a survey of 307 CHI reviewers, Nobarany et al. [61] found that “encouraging high-quality research, giving back to the research community, and finding out about new research” were reviewers’ primary motivations for reviewing. Our post-survey responses echoed these themes, describing science as a “community process” and emphasizing that “the fundamental purpose of peer review [...] is to provide expert, accountable, and independent judgment.” These findings suggest that, with the right culture and environment in place, reviewers would be motivated to conduct high-quality reviews. However, our findings suggest that these conditions are not in place, as evidenced by the high rates of noncompliance in both the Pangram analysis (Section 4.2.3) and post-survey results (Section 5.2.1).
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why does verification consistently lag behind AI generation?- Why is verification harder than generation across the research lifecycle?
- What makes proof writing and paper writing harder to verify than proof grading?
- How can automated review scale with the flood of AI-generated papers?
- What would it take for readers to inspect rather than assume authorship?
- What accountability structures should replace detection when AI automation increases in peer review?
- Can statistical filtering plus narrative generation fool academic peer review?
- What makes evaluative sophistication measurable in academic writing quality?
- Can LLMs evaluate their own observations without external feedback?
- Can proxy evaluation of ideas accurately predict their quality without implementation?
- Can researchers prevent their expectations from shaping LLM outputs?
- Do LLM judges with diverse personas resist individual biases better than single evaluators?
- Can LLMs reliably assess the quality of ideas they generate?
- What specific execution barriers do LLM ideas encounter most frequently?
- Why do LLM-generated ideas score higher novelty yet lower feasibility than expert ideas?
- Why do LLM research ideas lack diversity despite high average novelty?
- What role do multi-dimensional quality frameworks play in assessing arguments versus single-metric approaches?
- Can semantic clustering of stakeholders preserve meaningful evaluative diversity without manual curation?