LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv

Paper · arXiv 2601.17036 · Published January 19, 2026
Domain Specialization in LLMs

ArXiv recently prohibited the upload of unpublished review papers to its servers in the Computer Science domain, citing a high prevalence of LLM-generated content in these categories. However, this decision was not accompanied by quantitative evidence. In this work, we investigate this claim by measuring the proportion of LLM-generated content in review vs. nonreview research papers in recent years. Using two high-quality detection methods, we find a substantial increase in LLM-generated content across both review and non-review papers, with a higher prevalence in review papers. However, when considering the number of LLMgenerated papers published in each category, the estimates of non-review LLM-generated papers are almost six times higher. Furthermore, we find that this policy will affect papers in certain domains far more than others, with the CS subdiscipline Computers & Society potentially facing cuts of 50%. Our analysis provides an evidence-based framework for evaluating such policy decisions, and we release our code to facilitate future investigations at: https://gi thub.com/yanaiela/llm-review-arxiv.

Introduction. On October 31st, 2025, arXiv announced a policy change1 prohibiting the upload of review,2 survey, and position papers to their Computer Science (CS) servers (Boboris, 2025). According to the blog post, this decision was motivated by an observed increase in LLM-generated content in review papers in the CS subdomain over recent years. Since papers uploaded to arXiv undergo moderation by volunteer human experts, overburdened with other obligations, the administrators decided to restrict further approvals of unpublished review papers altogether. The blog post did not provide quantitative evidence or data to support this claim. In this paper, we address two central questions: Do review papers contain higher proportions of LLM-generated content than regular research papers? What downstream effects would be created by this policy? The proliferation of advanced AI writing assistants, particularly large language models (LLMs), has enabled researchers to incorporate these tools into their scientific workflows (Nejjar et al., 2025; Zhang et al., 2025; Eger et al., 2025; Si et al., 2025). In the extreme case, papers may be fully written by LLMs, including literature reviews and experimental design.3 While we have preliminary evidence for fully LLM-generated papers receiving scores sufficient for workshop publication (Yamada et al., 2025), a central concern is whether LLMgenerated papers maintain the same standards as human-written papers. The ease with which papers can be generated, and the increasing rates of paper submissions, especially about AI topics, provoke concerns about overwhelm for arXiv moderators. These are legitimate concerns, but the focus on review papers makes several assumptions that may or may not be accurate or helpful. First is the central assumption that review papers are more likely to be generated than non-review papers. Second is a set of epistemic assumptions about what constitutes a review paper and what purposes review papers serve in a research community, reflecting disciplinary biases (not shared across all CS subdisciplines) that elevate quantitative over qualitative methods and evidence. In this paper, we quantitatively evaluate this new policy. We evaluate the ratios of review papers that are LLM-generated vs. non-review papers, the temporal trends of such ratios, and how these trends vary across fields. Our main analysis classification and detection pipeline uses only paper titles and abstracts (not full text), reflecting the data used by arXiv to make the initial decision. We then run an additional analysis on a subset of papers for which we run the analysis on the full text, where we validate the consistency of our results. As such, we employ a two-stage methodology. First, we develop an LLM-based classifier to distinguish review papers from regular research papers, achieving high accuracy (92.0% F1) on manually annotated data. Second, we apply two LLM-generation detection methods that estimate the fraction of LLM-generated content in a corpus to measure LLM prevalence. We analyze a large dataset of arXiv papers from different domains to measure these rates. Our findings indicate that: (1) review papers contain higher rates of LLM-generated content than non-review papers in CS, with adjusted cohortlevel estimates of 21.4% for review papers versus 14.0% for non-review papers; (2) LLM-generated content has increased dramatically in both categories following ChatGPT’s release in November 2022, with review papers showing an increase from 12.9% in 2023 to 28.2% in 2025 and non-review papers from 6.2% to 18.9%; and (3) this pattern persists across all examined disciplines (Computer Science, Physics, Mathematics, Statistics) and shows an accelerating trend since 2023. Crucially, when considering the number of papers published for each type, we estimate that in CS, there are almost six times more non-review papers generated by LLMs than review papers (Figure 1). This result throws into doubt arXiv’s decision to restrict the upload of unpublished review papers to their CS servers. We also find that the policy will affect certain disciplines and topics differently; topics like education and safety have much higher rates of review papers, and Computers & Society papers could be censored by 50%, while Computer Vision will only lose 3% of its papers. We hope this work provides an evidence-based perspective on arXiv’s policy decision while fostering broader discussion about the long-term implications of LLMgenerated content in scientific publishing.

Related work. A Brief History of ArXiv. Arxiv originally started as an automated email server before converting into a website. As its usage and submissions grew, arXiv faced challenges related to software scale and content moderation (Han, 2025).4 While arXiv has historically relied on volunteer moderators to ensure basic quality standards, the scale and nature of submissions have evolved significantly, and as such, automatic classifiers were introduced to assist with moderation. The October 2025 policy restricting review papers in CS represents a notable departure from arXiv’s traditional open policy.5 LLM-Generated Text Detection. The detection of LLM-generated text has become an active research area following the release of powerful language models (Brown et al., 2020; Achiam et al., 2023). Early approaches relied on statistical features such as perplexity and token probability distributions (Mitchell et al., 2023). More recent work has explored zero-shot detection methods (Mitchell et al., 2023), watermarking techniques (Kirchenbauer et al., 2023), and distributional estimation approaches that measure the fraction of AI content in document collections rather than classifying individual texts (Sadasivan et al., 2023). Beyond the design of accurate detection models, there is the question about the right task definition. For instance, Ghosal et al. (2023) argues that binary AI-versushuman classifications are insufficient, proposing instead to estimate the percentage of LLM-generated LLMs in Academic Writing. Recent studies have documented the increasing use of LLM writing assistants and agents in academic contexts (Liang et al., 2024a). Concerns have been raised about the impact on scientific integrity, originality, and the peer review process (Lin et al., 2025; Latona et al., 2024). Several instances of fully or largely LLM-generated papers have been identified in published literature (Dupré, 2023), raising questions about quality control mechanisms. Some work has examined disciplinary differences in AI adoption, finding that fields with greater exposure to AI technology show earlier and higher adoption rates (Liao et al., 2025). However, to our knowledge, no prior work has systematically compared generated content between review papers and regular research papers or provided quantitative evidence for preprint server policies.

Method. We focus on estimating population-level trends (e.g., the proportion of LLM-generated content among all review papers posted on arXiv in 2025) and comparing these trends across groups (e.g., review vs. non-review, 2023 vs. 2025). While this is a challenging task, and we do not expect perfect accuracy, we can estimate and compare aggregate statistics over populations.6 We employ two complementary detection methods: (1) the Alpha estimator (Liang et al., 2024a), which provides a group-level estimate of the fraction of LLM-generated content based on the occurrence of adjectives used in a document, and (2) a commercial LLM-detector called Pangram (Emi The arXiv policy relies on distinguishing between “review” vs. non-review papers. But what is a “review” paper? While this question might seem obvious, such taxonomies are neither universal nor neutral; different fields can have different epistemic cultures and “machineries of knowledge construction” (Cetina, 1999), and interdisciplinary work often resists such categorizations (Clement, 2016). ArXiv’s CS corpus spans communities with different epistemic cultures and publication norms, from established subfields like Programming Languages (CS.PL) to interdisciplinary areas like Computation & Language (CS.CL) and Computers & Society (CS.CY). The latter two include papers from the digital humanities (DH), computational social science, and fairness, accountability, and transparency research (FAccT), where community norms are known to clash (Laufer et al., 2022). While we do not know exactly how arXiv moderators categorize review papers, we approximate what such a process might look like.9 What are Review Papers? Review papers, which we consider to include survey papers, often aim to synthesize and summarize a body of literature, identify trends over time, provide recommendations and guidelines for best practices, and introduce newcomers to a field. Structurally, review papers tend to emphasize literature coverage over novel experimental contributions, though high-quality reviews require substantial domain expertise. Some reviews offer novel insights through meta-analysis, and many different kinds of re- Review vs. Non-Review Dataset. To test the feasibility of the task and to evaluate the performance of our review paper classifier, we create a manually annotated dataset sampled from the arXiv Computer Science category across different years. We hand-annotate a set of 200 papers as review or nonreview papers. These papers were sampled such that 100 of them had keywords related to reviews10 We apply the review paper classifier to a large corpus of arXiv papers, classifying them into review and non-review papers. We then apply the AI content detector to both groups12 in each category and

Discussion. Our findings provide quantitative evidence that review papers on arXiv contain higher proportions of LLM-generated content than non-review papers, seemings to support arXiv’s stated rationale for their policy change. However, further analysis of the results reveals crucial nuances and unintended effects that warrant reconsideration of this policy.

Magnitude of the Difference. Over the past three years, we estimate the use of LLMs in CS papers to be 21.4% in review papers compared to 14.0% in non-review papers. Considering the peryear analysis, this gap increases considerably (e.g., a 6.7% difference in 2023 to 9.3% in 2025 using the Alpha estimate, and 4.7% in 2023 to 20.0% using Pangram). Crucially, however, when considering the number of papers from each type, the worry about LLM usage in reviews fades completely. We estimate the number of non-review papers to be 141K compared to 12K review papers in CS in 2025. Thus, even though the difference in review percentages is substantial in 2025 (43.3% vs. 23.3% using Pangram), the number of review papers generated by LLMs is estimated to be 4,783 compared to 26,801 for non-review papers (Figures 1 & 10, and Tables 6-9 in Appendix D.). As such, the arXiv ban on unpublished review papers makes little sense, if the goal is to save moderators time and energy, as the overall burden from review papers is tiny in comparison to the very large number of non-review papers using AI in paper writing.

Disciplinary Variations Matter. The crossdomain analysis (Figures 2 and 5) reveals that AI adoption patterns vary considerably across domains. CS shows the highest prevalence, unsurprising given the domain’s focus on AI technology and early access to these tools. However, the review versus non-review paper gap persists across most examined domains, suggesting this is a general phenomenon rather than one specific to CS. The temporal patterns are particularly revealing. The sharp increase following ChatGPT’s release in late 2022 is evident across all domains, but the rate of adoption differs. CS shows earlier and faster adoption, while traditional domains like Mathematics show more gradual increases. This suggests that domain-specific norms and tool familiarity might influence AI adoption rates. Our subcategory analysis within CS (§4.3) further demonstrates that even within a single domain, AI adoption varies dramatically. Certain subfields (Cryptography and Security, AI) show substantially higher adoption rates than others (Computation and Language, Information Retrieval), and the reviewnon-review gap manifests differently across subcategories. This fine-grained heterogeneity suggests that blanket policies applied uniformly across all CS subcategories may be overly broad. The validation of these subcategory patterns The Future of Scientific Publishing. Our crossdisciplinary findings suggest that LLM adoption in academic writing is not a temporary phenomenon but an ongoing trend that varies across fields. As LLM writing tools improve and become more integrated into more research workflows, we may see convergence in adoption rates across disciplines. On the other hand, as academic communities develop more established norms (ethical or otherwise), we may see further divergence. Finally, these patterns may evolve as both LLMs and detection models improve and as authors develop more sophisticated usage strategies. Publishers and preprint servers will need to develop clear, well-thought-out, and reasonable policies that account for disciplinary differences while maintaining scientific integrity and considering the development and incorporation of AI-assisted tools.

Conclusion. Our study evaluates the evidence behind arXiv’s 2025 restriction on review papers in the CS domain. Across two independent LLM detectors (Alpha and Pangram) and an accurate review vs. non-review classifier, we find that review papers consistently contain higher proportions of LLM-generated content. However, the much larger volume of nonreview papers means that they account for the vast majority of LLM-generated manuscripts on arXiv, challenging the premise of a review-only ban. Temporal and disciplinary analyses reveal a sharp post-ChatGPT rise in AI-assisted writing that spans CS, Physics, Mathematics, and Statistics, with heterogeneous uptake across CS subfields. These patterns suggest that LLM-assisted writing is a broad, accelerating shift rather than a CS or review-specific anomaly. Blanket restrictions on review papers risk missing most LLM-generated content while disproportionately affecting communities where synthesis and qualitative work are core scholarship. We advocate for evidence-based moderation that combines transparent monitoring with clearer guidance on acceptable LLM assistance.

Limitations. Our study has several important limitations that should be considered when interpreting the results.

Detection Method Limitations. The Alpha estimator, while powerful for group-level estimation, relies on distributional assumptions about LLMgenerated text. The method may be sensitive to changes in the LLM used for generation, the data used to estimate alpha, or post-editing by human authors. Our pre-LLM baseline validation suggests the Alpha estimator has differential false positive rates for review and regular papers, which we address through type-specific Rogan-Gladen adjustment. However, this assumes the false positive pattern remains stable over time and does not interact with actual AI use. On the other hand, we detected no false positives with the Pangram detector.

Classification Accuracy. Our review vs. nonreview paper classifier achieves a high classification score (92.0% F1), but misclassifications could still bias our results. We validated the classifier on held-out data, but some borderline cases (e.g., tutorial papers, hybrid papers with both novel contributions and community norms) may be ambiguous. In general, our annotations reflect an attempt at a “typical” CS view of what a “review” paper might represent, as this is likely the stance of many arXiv moderators, rather than trying to represent the many epistemic views of different subfields.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How can we detect and account for LLM involvement in academic writing? How can AI systems reliably guide voters without introducing political bias? Can AI systems perform peer review as effectively as humans? Do restrictions on reviewer LLM use actually shape peer review behavior? What human oversight must AI research systems have? How do educators verify student capability when AI can produce indistinguishable work? How can we reduce inherent biases in LLM-based evaluation judges? Does AI assistance help or harm professional skill development?