AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot

Paper · arXiv 2604.13940 · Published April 15, 2026
Domain Specialization in LLMs

Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances in AI have led the community to consider its use in peer review, yet a key unresolved question is whether AI can generate technically sound reviews at real-world conference scale. Here we report the first large-scale field deployment of AIassisted peer review: every main-track submission at AAAI- 26 received one clearly identified AI review from a state-ofthe-art system. The system combined frontier models, tool use, and safeguards in a multi-stage process to generate reviews for all 22,977 full-review papers in less than a day. A large-scale survey of AAAI-26 authors and program committee members showed that participants not only found AI reviews useful, but actually preferred them to human reviews on key dimensions such as technical accuracy and research suggestions. We also introduce a novel benchmark and find that our system substantially outperforms a simple LLMgenerated review baseline at detecting a variety of scientific weaknesses. Together, these results show that state-of-the-art AI methods can already make meaningful contributions to scientific peer review at conference scale, opening a path toward the next generation of synergistic human-AI teaming for evaluating research.

Introduction. The scientific peer review process is under significant strain. The AAAI Conference on Artificial Intelligence, a major artificial intelligence (AI) research conference, received more than 30,000 initial submissions for 20261, up from approximately 15,000 for 2025. This dramatic growth is not unique to AAAI; submissions have grown rapidly for other venues too, such as Nature [1] and NeurIPS [2]. Unfortunately, despite this rapid growth, the peer review process has remained largely static, with a large cohort of human reviewers providing detailed reviews and ratings for papers, and a smaller group of senior researchers comparing those reviews and ratings to make final paper acceptance recommendations. While this established review process has long endured, the rising scale of submissions means we face increasingly overburdened reviewers with more papers assigned, the need to recruit an ever-wider pool of potentially less experienced reviewers, and increasingly compressed timelines. Maintaining the quality, consistency, and timeliness of peer review is thus increasingly challenging. For example, the scale of AAAI-26 submissions required the recruitment and oversight of over 28,000 Program Committee members, Senior Program Committee members, and Area Chairs, nearly three times the size of the committee in AAAI-25 [3]. At the same time, there have been rapid advances in stateof-the-art AI systems, particularly in mathematical, coding [4, 5], and other technical domains [6]. Autonomous AI scientists [7, 8, 9, 10, 2] now perform iterative feedback loops of automated paper writing, generating critiques, and revising the writing in response. Most relevant, a growing body of work is now investigating whether and how AI systems can be used to assist with scientific peer review and help alleviate the growing strain on the review process [11, 12, 13, 14, 15, 16, 17, 18, 19, 20]. Against the backdrop of increasing strain on human peer review, the reviewer population has started using AI reviewing against explicit guidance [21], while conference venues grapple with the question of how to effectively and meaningfully integrate AI into the review process in a way that is beneficial to the community. There have been synthetic, benchmark-based, and posthoc studies of AI-generated reviews and AI-assisted peer review on existing papers and reviews, including retrospective analyses of reviewers’ use of AI during peer review [12, 14, 16, 18, 17, 13, 22]. Evaluation remains challenging, however, because existing benchmarks and evaluation datasets measure only limited aspects of reviewing. These aspects include specific error types, evaluation of structured outputs as opposed to unstructured review text, or similarity of scores to human reviews, rather than end-toend review quality [23, 24, 25, 26].

Based on these encouraging findings, there have been a small number of live studies of AI systems providing limited assistance within conference workflows, notably author checklist assistance in NeurIPS 2024 [27] and feedback to reviewers in ICLR 2025 [28]. However, neither of these studies deployed official AI-generated reviews on live submissions. Prior to AAAI-26, there had been no conferencewide live study of AI-generated reviews deployed on real submissions at a major conference. Thus, despite substantial recent progress, a key question remained: could stateof-the-art AI systems generate technically meaningful and practically useful reviews in a live peer-review process at conference scale? The AAAI-26 AI Review Pilot Program was the first fullscale live study of AI-generated reviews on real submissions at a major conference. Every paper that entered the full review phase (22,977 in total) in the main track at AAAI-26 received one clearly labeled AI review, generated by a stateof-the-art custom-developed AI review system. Consistent with prior results, we found that simply asking off-theshelf LLMs to review papers does not lead to high-quality reviews [14, 16, 29]. Recent work has therefore explored more structured review systems based on deeper multi-stage reasoning, hierarchical question decomposition, and multimodal workflows [30, 31, 32]. In response, we developed a novel, multi-stage, multi-tool, LLM-based review pipeline that does lead to very high-quality reviews. AAAI-26 used a double-blind review process, so reviewers and authors were anonymized to one another during evaluation, but both reviewers and authors could identify the AI review. The AI review was added during Phase 1 of the two-phase review process, alongside at least two human reviews. The AI review system did not include any scores or recommendations, and no human reviewers were replaced in the process. Instead, the AI reviews were intended to provide additional input to the peer-review process [33].

Related work. Based on these encouraging findings, there have been a small number of live studies of AI systems providing limited assistance within conference workflows, notably author checklist assistance in NeurIPS 2024 [27] and feedback to reviewers in ICLR 2025 [28]. However, neither of these studies deployed official AI-generated reviews on live submissions. Prior to AAAI-26, there had been no conferencewide live study of AI-generated reviews deployed on real submissions at a major conference. Thus, despite substantial recent progress, a key question remained: could stateof-the-art AI systems generate technically meaningful and practically useful reviews in a live peer-review process at conference scale? The AAAI-26 AI Review Pilot Program was the first fullscale live study of AI-generated reviews on real submissions at a major conference. Every paper that entered the full review phase (22,977 in total) in the main track at AAAI-26 received one clearly labeled AI review, generated by a stateof-the-art custom-developed AI review system. Consistent with prior results, we found that simply asking off-theshelf LLMs to review papers does not lead to high-quality reviews [14, 16, 29]. Recent work has therefore explored more structured review systems based on deeper multi-stage reasoning, hierarchical question decomposition, and multimodal workflows [30, 31, 32]. In response, we developed a novel, multi-stage, multi-tool, LLM-based review pipeline that does lead to very high-quality reviews. AAAI-26 used a double-blind review process, so reviewers and authors were anonymized to one another during evaluation, but both reviewers and authors could identify the AI review. The AI review was added during Phase 1 of the two-phase review process, alongside at least two human reviews. The AI review system did not include any scores or recommendations, and no human reviewers were replaced in the process. Instead, the AI reviews were intended to provide additional input to the peer-review process [33]. Senior Program Committee members (SPCs) and Area Chairs (ACs) (who are responsible for making paper recommendations and normalizing within their batch), were able to view the AI reviews along with the human reviews and use them to help make their decisions about whether to promote papers to Phase 2 of the review process. Papers that were promoted to Phase 2 received additional human reviews, and the authors had a chance to respond to all reviews, including the AI review, before the reviewers, SPCs, and ACs discussed the papers and made their final decisions in light of all reviews, including the AI review. An optional survey was sent to authors, reviewers, SPCs, and ACs to assess both human and AI reviews on a variety of criteria.

Method. 2 The AAAI-26 AI Review System The AAAI-26 AI Review System integrates learnings from prior studies of AI-generated reviews [12, 16, 29, 7]. A key design goal for the system was to ensure that the reviews considered scientific accuracy of all forms — including mathematical and algorithmic correctness, sufficiency of the evaluation methods, and positioning of the work in the context of the previous state-of-the-art. Previous studies have shown that prompting LLMs with different ‘personas’ [24, 18, 17] or criterion-specific prompts [14], rather than asking them to directly produce full reviews, can improve identification of specific types of scientific errors. Recent systems have also explored hierarchical question decomposition, deeper staged reasoning, and multimodal agent designs with shared memory for paper review [31, 30, 32]. The AAAI-26 AI Review System thus consists of five core scientific review stages intended to identify errors in: (1) story, (2) presentation, (3) evaluations, (4) correctness, and (5) significance. The review system takes each PDF paper as input and generates a textual review with markdown notation [35] for math and tables, such that it can be rendered on the review interface.

Fig. 1 shows the multi-stage AAAI-26 AI Review System that we built based on the desiderata and insights above. After preprocessing, subsequent stages of the review system include both PDF and markdown versions of the paper, along with a system prompt to provide context for when to pay particular attention to one version over the other. Each review stage includes both a stage-specific prompt as well as the prompts and results from all previous stages. The evaluations and correctness stages include a Python code interpreter made available to the LLM to allow it to check for errors by testing out math and code snippets. The significance stage includes a web search tool to assist with literature search — with specific instructions to restrict references to published work at relevant venues. After all targeted stages, the system generates an initial review and then revises it through a self-critique stage. The initial review generation and revision prompts include specific instructions to ensure that each review contains the following structural elements: (1) the title of the paper, (2) a brief synopsis of the paper, (3) a summary of the review, (4) a detailed list of strengths, (5) a detailed list of weaknesses, and (6) a list of references cited in the review, in APA citation format. Prompt details for each stage are provided in Section A. We implemented a quality-checking workflow to identify potential issues in the generated reviews, similar to previous work on “peer reviews of peer reviews” [36]. Details on the quality-checking workflow and additional checks for citation hallucinations are provided in Section A.

Discussion. AI reviews created additional cognitive work for authors and other reviewers (Excessive Verbosity and Cognitive Overload). Such verbosity and nitpicking highlights the downside of the thoroughness found in AI reviews. Finally, the feedback also discussed Factual Errors and Misreadings found in AI reviews, indicative of a continuing gap in LLMs understanding scientific content compared to an academic (human) established within their field. This ties into the the fifth most discussed theme (Shallow Contextual and Domain Understanding), which specifically points out how the AI reviews struggled to provide feedback that was appropriate to a particular research domain.

9 out of the 33 themes we found within the written feedback were general opinions about the use of AI in academic review process that were not specific to the AAAI-26 AI Review Pilot. Representative free-form survey responses illustrating these patterns are provided in Section A. Respondents also spoke positively about using AI as a presubmission assistance tool and a means of generating metalevel summaries [37]. They highlighted its potential to scale peer review and support overburdened reviewers, while acknowledging that these systems are poised for rapid future improvement. Not all of these opinionated themes were positive. Respondents also emphasized that AI reviews had the potential to mislead reviewers and other decision-makers in the review process. There were also concerns that authors might optimize papers for AI preferences rather than scientific quality, and that reliance on these tools could lead to a long-term decline in reviewing skill. Adding to this, many respondents voiced principled objections, arguing that the use of AI undermines the trust, human effort, and essential value of the peer review process.

Conclusion. The AAAI-26 AI Review Pilot Program demonstrated that AI-generated peer reviews are operationally feasible at conference scale, and are capable of generating reviews that are helpful to the authors and reviewers. The large-scale survey of AAAI-26 authors, reviewers, senior program committee members, and area chairs found that participants broadly found AI reviews useful and preferred them to human reviews on key dimensions such as technical accuracy and research suggestions, but also identified some limitations and areas for improvement including technical errors in reading some equations and tables, difficulty in prioritizing the significance of issues, and producing reviews that were longer than readers preferred. The quantitative and qualitative analyses combined indicate the complementary strengths of AI systems and human reviewers, and suggest that future work should explore how best to integrate the two to leverage their strengths.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? Can AI systems discover fundamental improvements to their own architectures? How do hallucinated citations emerge in AI scholarly output? How can humans maintain effective oversight as AI systems scale? How can evaluations be made robust against model reward hacking? What explains the gap between benchmark scores and true reasoning capability? What human oversight must AI research systems have? Can AI research automation sustain progress through accelerating feedback loops? Do restrictions on reviewer LLM use actually shape peer review behavior?