The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing

Paper · arXiv 2609.07713 · Published September 7, 2026
Domain Specialization in LLMs

Generative and agentic AI are reshaping both the production and evaluation of scientific research. These developments are often studied separately, as questions of how AI can produce research and how AI can review it. We argue that this separation misses an increasingly important feature of scholarly publishing: changes on one side alter the incentives, constraints, and behavior of the other. We synthesize 230 scholarly publications and institutional records using a taxonomy of six connected dynamics: production scaling, evaluation automation, evaluation manipulation, defense mechanisms and policy responses, evasion and side effects, and long-horizon ecosystem feedback. The literature shows an emerging progression in which cheaper and faster research production increases pressure on evaluation, AI-mediated evaluation becomes more scalable and repeatable, participants can exploit evaluator regularities, and institutions respond with technical safeguards and policy controls. These responses can in turn induce evasion, redistribute errors and workload, and shape the scholarly records reused by future research and evaluation systems. Evidence is strongest for production and evaluation at scale, reproducible manipulation, and institutional response, while post-policy adaptation and artifact-level long-horizon feedback remain less directly observed. This systems view shifts attention from isolated AI capabilities toward how scholarly actors and AI systems adapt to one another over time.

Introduction. Generative and agentic AI are increasingly entering both sides of scientific research and evaluation. On the production side, researchers use LLM-based tools throughout the research cycle, from exploring ideas and reviewing relevant literature to conducting experiments and preparing manuscripts (Baek et al., 2025; Liang et al., 2025; Lu et al., 2024; Yamada et al., 2025; Tang et al., 2026b; Kong et al., 2026). On the evaluation side, generative AI is increasingly used across peer review, extending from manuscript summarization and critique generation to evaluation and publication decisionmaking (Liang et al., 2024b; Nguyen and Ahmadi, 2026; Thakkar et al., 2026; Biswas et al., 2026). These developments are commonly studied through two corresponding questions: how AI can support or automate scientific research, and how AI can perform or assist scholarly evaluation. While this separation is useful for studying the capabilities of autonomous research systems and AI reviewers in isolation, it is incomplete once the two sides operate within the same publishing process. AI-mediated production changes the volume, speed, and form of what enters evaluation, while AI-mediated evaluation creates new signals and regularities to which authors can respond. The relevant challenge therefore lies not only in the capabilities of either side, but also in how changes on one side reshape the conditions, costs, and incentives faced by the other.

The most immediate form of this coupling is a capacity imbalance. As AI reduces the time and marginal cost required to conduct, write, and revise research, scientific output and submission volume can grow faster than available human reviewing capacity, increasing pressure to scale evaluation as well. Evidence for this dynamic is already emerging from empirical studies and real-world deployments. Observational studies associate LLM adoption with shorter research and revision cycles, increased scientific output, rising submission volume, and growing pressure on limited reviewing capacity (Kusumegi et al., 2025; Qi et al., 2025; Filimonovic et al., 2025; Gartenberg et al., 2026). Increasingly capable agentic AI systems extend this scaling from assistance with individual tasks toward integrated idea-to-paper workflows, further reducing the time and marginal cost required to produce complete research artifacts (Kong et al., 2026; Tang et al., 2026b; Song et al., 2026; Yuan et al., 2026). Evaluation is beginning to scale alongside this growth: AI is now used both to assist individual reviewers (Thakkar et al., 2026) and within official conference workflows, as illustrated by the large-scale AAAI-26 AI-review pilot (Biswas et al., 2026). At the same time, machine-mediated evaluation makes parts of the evaluation process more repeatable and therefore easier for authors to adapt to. Hidden instructions targeting AI-assisted reviewers have been identified in public manuscripts (Lin, 2026), while controlled studies show that reviewer judgments can be shifted without corresponding improvements in the underlying scientific evidence and that repeated evaluator feedback can be used to optimize manuscripts toward more favorable assessments (Yang et al., 2026c; Zhou et al., 2025; Li et al., 2026c). Venues have, in turn, introduced AI-specific review policies, submission controls, detection mechanisms, and integrity safeguards (ICLR 2026 Program Chairs, 2025a; Agarwal et al., 2026; NeurIPS 2026 Communication Chairs, 2026; Li et al., 2024b; Fichtl et al., 2026). Taken together, these developments form an emerging sequence of scaling, adaptation, and institutional response rather than two independent trends in AI adoption.

Figure 1 summarizes this coupled sequence and the scholarly actors through which it operates.

To better characterize this adaptive process and provide a coherent structure for the survey, we develop a descriptive taxonomy of six linked dynamics: (1) production scaling, (2) evaluation automation, (3) evaluation manipulation, (4) defense mechanisms and policy responses, (5) evasion and side effects, and (6) long-horizon ecosystem feedback. The taxonomy is organized around the principal response relations among actors in the scientific ecosystem rather than around technologies alone. Production scaling changes the demand placed on evaluation; evaluation automation makes parts of the evaluation process increasingly repeatable and therefore more susceptible to manipulation and optimization; venues respond through technical and institutional defenses; and those interventions can induce evasion or shift costs and risks to other actors. The resulting papers, reviews, and decisions can also persist and influence future research and evaluation systems. These dynamics are therefore relation-oriented and non-exclusive: each reported finding, deployment, or institutional action is assigned a primary dynamic according to the scholarly function it directly observes, tests, or implements.

Related work. Existing surveys provide valuable accounts of AI-assisted scientific production and peer review, typically organizing the literature around tasks, system components, or stages of the research and review workflow (Kuznetsov et al., 2024; Zhuang et al., 2025; Wu et al., 2026c; Nguyen and Ahmadi, 2026; Kong et al., 2026). Recent surveys have also highlighted the verification gap in autonomous research agents and discussed governance and human oversight in AI-assisted peer review (Ding et al., 2026b; Mann et al., 2025; Wei et al., 2025). Other recent work has also begun to connect review with rebuttal, revision, and further experimentation (Weng et al., 2025; Han et al., 2026; Ma et al., 2026b; Wu, 2026). These perspectives explain what individual systems do and how outputs move through a workflow, but they do not fully capture how one actor’s use of AI changes another actor’s signals, costs, and feasible responses. Production scaling can alter evaluation demand, machine-mediated evaluation can create new targets for optimization, and institutional interventions can change the incentives for subsequent adaptation and evasion. What is therefore missing is not another inventory of AI tools, but a relational account of how these developments interact across actors in the scientific ecosystem. We synthesize these relations as a coupled process that we term the AI paper-review arms race. The term serves as a descriptive lens for documented or testable patterns of adaptation and counter-adaptation across actors in the scientific ecosystem.

Method. This section formalizes the taxonomy used throughout the survey. We first define the system boundary and the institutional roles through which AI operates within the scientific ecosystem, then describe the six dynamics and their principal response order, and finally specify the literature-mapping and source-selection procedures. The taxonomy is descriptive and relation-oriented, organizing the literature around scholarly functions and the response relations examined across studies.

This survey examines the coupled and recursive effects of emerging AI tools across the scientific ecosystem. The scope begins with research production, including idea development, experimentation, manuscript preparation, revision, and rebuttal, and extends to peer review, publication decisions, and the policies, tools, and safeguards introduced by conferences and publishers. We also consider how these stages interact, as changes in research production can alter evaluation practices, while evaluation and institutional responses can in turn shape subsequent author behavior. Beyond these immediate interactions, we examine their longer-term consequences, including how papers, reviews, decisions, corrections, and citations can be reused by retrieval systems, scientific agents, training pipelines, and AI-based evaluators, thereby influencing future rounds of research and evaluation. The survey is organized around the progression of these interactions and responses. Section 3 examines Under this scope, we distinguish several functional roles rather than a fixed set of mutually exclusive actors. Authors and AI-enabled research systems produce and revise scientific work and respond to evaluation, while reviewers, editors, and committees assess that work and contribute to publication decisions. At the institutional level, venues, publishers, and platform operators shape the rules and technical conditions under which research and evaluation take place, including review policies, evaluation rubrics, and safeguards. AI-enabled systems may operate within these constraints and may also reuse papers, reviews, decisions, and other scholarly records in subsequent research or evaluation. These roles can overlap, and the same actor or system may occupy multiple roles. The taxonomy therefore distinguishes scholarly functions rather than fixed actor identities, allowing us to trace how changes in one part of the ecosystem alter the conditions and responses of others.

We use the paper-review arms race as an analytic lens for sequences of adaptive response, not as a synonym for any use of AI in research or review. Operationally, an arms-race episode involves an actor taking an action that changes a target or signal relevant to another actor, a counter-response by that actor or an institution, and a resulting shift in incentives, costs, or capabilities that can motivate further adaptation. A single instance of AI-assisted production, evaluation, or defense can therefore belong to the taxonomy without constituting an arms-race episode unless it participates in such a response relation. Because complete sequences are rarely observed within a single study, connections synthesized across separate bodies of literature are presented as research questions rather than as directly observed causal sequences.

Our taxonomy comprises six process categories that describe how AI changes scientific research and evaluation and how different parts of the ecosystem respond to those changes. We refer to these categories as dynamics because they capture processes of change and response rather than static system components. Table 1 situates this relation-oriented organization relative to recent surveys organized around research and review tasks or workflows. Figure 2 provides a detailed hierarchy of the six dynamics, their recurring mechanisms, and representative studies.

Discussion. 9 Cross-cutting findings and research agenda 9.1 Trustworthy evaluation remains the bottleneck Generative and agentic systems can increasingly reduce the time and effort required to produce, revise, and evaluate research artifacts. The corresponding work of verifying scientific claims, identifying consequential flaws, assessing novelty, and making accountable decisions remains much harder to scale. Evidence from journals and conferences already connects growing production volume with pressure on editorial and review capacity, while automated review systems can generate feedback at substantially lower marginal cost (Gartenberg et al., 2026; Knight, 2026; Biswas et al., 2026).

The resulting bottleneck is therefore not simply review generation, but trustworthy scientific evaluation. Faster review generation is useful only when the resulting judgments remain grounded and when the human effort required for verification, disagreement resolution, and oversight does not grow at the same rate. Future deployments should therefore measure evaluation capacity in terms of both throughput and the human work required to maintain reliable decisions.

9.2 Adaptation weakens static evaluation Evaluation should therefore test not only performance on a fixed distribution, but also whether judgments remain reliable under adaptation. Relevant questions include whether semantically similar papers receive consistent evaluations, whether important scientific weaknesses remain influential after presentation changes, and whether robustness persists when participants receive feedback and can change their strategies (Dycke and Gurevych, 2026; Jiang et al., 2026a). The central issue is whether an evaluator continues to distinguish stronger scientific work from more effectively optimized presentation once its behavior becomes part of the environment.

The evidence also shows that defensive measures rarely remove risk without changing where it appears. Detection and verification can make particular forms of AI use or manipulation more costly, but they can also motivate evasion. More aggressive enforcement can introduce false positives, while AI-assisted evaluation can create additional verification work, confidentiality concerns, and unequal effects across participants (Liang et al., 2023; Saha et al., 2026; Boonsuk, 2026; Vasu et al., 2026; Chen et al., 2025).

These directions shift the research agenda from evaluating isolated AI components toward studying the response relations that connect production, evaluation, manipulation, defense, evasion, and longhorizon feedback. Progress will depend especially on evidence that follows these interactions in real scholarly settings and over time.

Conclusion.

Limitations. Finally, the paper-review arms-race framing emphasizes adaptation and counter-adaptation between scholarly actors. It is intended as a lens for organizing these interactions rather than as a claim that all AI-assisted research, reviewing, revision, or institutional change is adversarial. Many uses of AI can improve access, efficiency, communication, and scientific quality, and the relevance of the framing depends on whether one actor’s behavior meaningfully changes the incentives or responses of another.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? What human oversight must AI research systems have? How should human-AI contributions be measured, disclosed, and verified? Does AI-assisted research sacrifice exploration breadth for productivity gains? Do restrictions on reviewer LLM use actually shape peer review behavior? How do hallucinated citations emerge in AI scholarly output?