Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews

Paper · arXiv 2403.07183 · Published March 11, 2024
Domain Specialization in LLMs

We present an approach for estimating the fraction of text in a large corpus which is likely to be substantially modified or produced by a large language model (LLM). Our maximum likelihood model leverages expert-written and AI-generated reference texts to accurately and efficiently examine real-world LLM-use at the corpus level. We apply this approach to a case study of scientific peer review in AI conferences that took place after the release of ChatGPT: ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023. Our results suggest that between 6.5% and 16.9% of text submitted as peer reviews to these conferences could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates. The circumstances in which generated text occurs offer insight into user behavior: the estimated fraction of LLM-generated text is higher in reviews which report lower confidence, were submitted close to the deadline, and from reviewers who are less likely to respond to author rebuttals. We also observe corpus-level trends in generated text which may be too subtle to detect at the individual level, and discuss the implications of such trends on peer review. We call for future interdisciplinary work to examine how LLM use is changing our information and knowledge practices.

Introduction. Despite the fact that generated text may be indistinguishable on a case-by-case basis from content written by humans, studies of LLM-use at scale find corpus-level trends which contrast with at-scale human behavior. For example, the increased consistency of LLM output can amplify biases at the corpus-level in a way that is too subtle to grasp by examining individual cases of use. Bommasani et al. find that the “monocultural” use of a single algorithm for hiring decisions can lead to “outcome homogenization” of who gets hired—an effect which could not be detected by evaluating hiring decisions one-by-one. Cao et al. find that prompts to ChatGPT in certain languages can reduce the variance in model responses, “flattening out cultural differences and biasing them towards American culture”; a subtle yet persistent effect that would be impossible to detect at an individual level. These studies rely on experiments and simulations to demonstrate the importance of analyzing and evaluating LLM output at an aggregate level. As LLM-generated content spreads to increasingly highstakes information ecosystems, there is an urgent need for efficient methods which allow for comparable evaluations on real-world datasets which contain uncertain amounts of AI-generated text.

We propose a new framework to efficiently monitor AImodified content in an information ecosystem: distributional GPT quantification (Figure 2). In contrast with instance-level detection, this framework focuses on population-level estimates (Section § 3.1). We demonstrate how to estimate the proportion of content in a given corpus that has been generated or significantly modified by AI, without the need to perform inference on any individual instance (Section § 3.2). Framing the challenge as a parametric inference problem, we combine reference text which is known to be human-written or AI-generated with a maximum likelihood estimation (MLE) of text from uncertain origins (Section § 3.3). Our approach is more than 10 million times (i.e., 7 orders of magnitude) more computationally efficient than state-of-the-art AI text detection methods (Table 20), while still outperforming them by reducing the in-distribution estimation error by a factor of 3.4, and the out-of-distribution estimation error by a factor of 4.6 (Section § 4.2,4.3).

Inspired by empirical evidence that the usage frequency of these specific adjectives like “commendable” suddenly spikes in the most recent ICLR reviews (Figure 1), we run systematic validation experiments to show that these adjectives occur disproportionately more frequently in AIgenerated texts than in human-written reviews (Supp. Table 2,3, Supp. Figure 12,13). These adjectives allow us to parameterize our compound probability distribution framework (Section § 3.5), thereby producing more empirically stable and pronounced results (Section § 4.2, Figure 3). However, we also demonstrate that similar results can be achieved with adverbs, verbs, and non-technical nouns (Ap- We demonstrate this approach through an in-depth case study of texts submitted as reviews to several top AI conferences, including ICLR, NeurIPS, EMNLP, and CoRL (Section § 4.1, Table 1) as well as through reviews submitted to the Nature family journals (Section § 4.4). We find evidence that a small but significant fraction of reviews written for AI conferences after the release of ChatGPT could be substantially modified by AI beyond simple grammar and spell checking (Section § 4.5,4.6, Figure 4,5,6). In contrast, we do not detect this change in reviews in Nature family journals (Figure 4), and we did not observe a similar trend of Figure 1 (Section § 4.4). Finally, we show several ways to measure the implications of generated text in this information ecosystem (Section § 4.7). First, we explore the circumstances in AI-generated text appears more frequently, and second, we demonstrate how AI-generated text appears to differ from expert-written reviews at the corpus level (See summary in Box 1).

Throughout this paper, we refer to texts written by human experts as “peer reviews” and texts produced by LLMs as “generated texts“. We do not intend to make an ontological claim as to whether generated texts constitute peer reviews; any such implication through our word choice is unintended.

In summary, our contributions are as follows:

  1. We propose a simple and effective method for estimating the fraction of text in a large corpus that has been substantially modified or generated by AI (Section § 3). The method uses historical data known to be human expert or AI-generated (Section § 3.4), and leverages this data to compute an estimate for the fraction of AI-generated text in the target corpus via a maximum likelihood approach (Section § 3.5).

Related work. Zero-shot LLM detection. Many approaches to LLM detection aim to detect AI-generated text at the level of individual documents. Zero-shot detection or “model selfdetection” represents a major approach family, utilizing the heuristic that text generated by an LLM will exhibit distinctive probabilistic or geometric characteristics within the very model that produced it. Early methods for LLM detec- Training-based LLM detection. An alternative LLM detection approach is to fine-tune a pretrained model on datasets with both human and AI-generated text examples in order to distinguish between the two types of text, bypassing the need for original model access. Earlier studies have used classifiers to detect synthetic text in peer review corpora (Bhagat & Hovy, 2013), media outlets (Zellers et al., 2019), and other contexts (Bakhtin et al., 2019; Uchendu et al., 2020). More recently, GPT-Sentinel (Chen et al., 2023) train the RoBERTa (Liu et al., 2019) and T5 (Raffel et al., 2020) classifiers on the constructed dataset OpenGPT- Text. GPT-Pat (Yu et al., 2023) train a twin neural network to compute the similarity between original and re-decoded texts. Li et al. (2023) build a wild testbed by gathering texts from various human writings and deepfake texts generated by different LLMs. Notably, the application of contrastive and adversarial learning techniques has enhanced classifier robustness (Liu et al., 2022; Bhattacharjee et al., 2023; Hu et al., 2023a). However, the recent development of several publicly available tools aimed at mitigating the risks associated with AI-generated content has sparked a debate about their effectiveness and reliability (OpenAI, 2019; Jawahar et al., 2020; Fagni et al., 2021; Ippolito et al., 2019; Mitchell et al., 2023b; Gehrmann et al., 2019; Heikkil ̈a, 2022; Crothers et al., 2022; Solaiman et al., 2019a). This discussion gained further attention with OpenAI’s 2023 decision to discontinue its AI-generated text classifier due to its “low rate of accuracy” (Kirchner et al., 2023; Kelly, 2023).

A major empirical challenge for training-based methods is their tendency to overfit to both training data and language models. Therefore, many classifiers show vulnerability to adversarial attacks (Wolff, 2020) and display bias towards writers of non-dominant language varieties (Liang et al., 2023a). The theoretical possibility of achieving accurate instance-level detection has also been questioned by researchers, with debates exploring whether reliably distinguishing AI-generated content from human-created text on an individual basis is fundamentally impossible (Weber- Wulff et al., 2023; Sadasivan et al., 2023; Chakraborty et al., 2023).

Method. 3.1. Notation & Problem Statement Let x represent a document or sentence, and let t be a token. We write t ∈x if the token t occurs in the document x. We will use the notation X to refer to a corpus (i.e., a collection of individual documents or sentences x) and V to refer to a vocabulary (i.e., a collection of tokens t). In all of our experiments in the main body of the paper, we take the vocabulary V to be the set of all adjectives. Experiments comparing against these other possibilities such as adverbs, verbs, nouns can be found in the Appendix D.5,D.6,D.7. That is, all of our calculations depend only on the adjectives contained in each document. We found this vocabulary choice to exhibit greater stability than using other parts of speech such as adverbs, verbs, nouns, or all possible tokens. We removed technical terms by excluding the set of all technical keywords as self-reported by the authors during abstract submission on OpenReview.

Let P and Q denote the probability distribution of documents written by scientists and generated by AI, respectively. Given a document x, we will use P(x) (resp. Q(x)) to denote the likelihood of x under P (resp. Q). We assume that the documents in the target corpus are generated from the mixture distribution and the goal is to estimate the fraction α which are AIgenerated.

3.2. Overview of Our Statistical Estimation Approach LLM detectors are known to have unstable performance (Section § 4.3). Thus, rather than trying to classify each document in the corpus and directly count the number of occurrences in this manner, we take a maximum likelihood approach. Our method has three components: training data generation, document probability distribution estimation, and computing the final estimate of the fraction of text that has been substantially modified or generated by AI. The method is summarized graphically in Figure 2. A nongraphical summary is as follows:

  1. Collect the writing instructions given to (human) authors for the original corpusin our case, peer review instructions. Give these instructions as prompts into an LLM to generate a corresponding corpus of AIgenerated documents (Section § 3.4). 2. Using the human and AI document corpora, estimate the reference token usage distributions P and Q (Section § 3.5). 3. Verify the method’s performance on synthetic target corpora where the correct proportion of AI-generated documents is known (Section § 3.6). 4. Based on these estimates for P and Q, use MLE to estimate the fraction α of AI-generated or modified documents in the target corpus (Section § 3.3).

The following sections present each of these steps in more detail.

If P and Q are known, we can then estimate α via maximum likelihood estimation (MLE) on (2). This is the final step in our method. It remains to construct accurate estimates for P and Q.

3.4. Generating the Training Data We require access to historical data for estimating P and Q. Specifically, we assume that we have access to a collection of reviews which are known to contain only human-authored text, along with the associated review questions and the reviewed papers. We refer to the collection of such documents as the human corpus.

To generate the AI corpus, we prompt the LLM to generate a review given a paper. The texts output by the LLM are then collected into the AI corpus. Empirically, we found that our 3.5. Estimating P and Q from Data The space of all possible documents is too large to estimate P(x), Q(x) directly. Thus, we make some simplifying assumptions on the document generation process to make the estimation tractable.

We represent each document xi as a list of occurrences (i.e., a set) of tokens rather than a list of token counts. While longer documents will tend to have more unique tokens (and thus a lower likelihood in this model), the number of additional unique tokens is likely sublinear in the document length, leading to a less exaggerated down-weighting of longer documents.1 The occurrence probabilities for the human document distribution can be estimated by where X is the corpus of human-written documents. The estimate ˆq(t) can be defined similarly for the AI distribution.

Discussion. In this work, we propose a method for estimating the fraction of documents in a large corpus which were generated primarily using AI tools. The method makes use of historical documents. The prompts from this historical corpus are then fed into an LLM (or LLMs) to produce a corresponding corpus of AI-generated texts. The written and AI-generated corpora are then used to estimate the distributions of AIgenerated vs. written texts in a mixed corpus. Next, these estimated document distributions are used to compute the likelihood of the target corpus, and the estimate for α is produced by maximizing the likelihood. We also provide specific methods for estimating the text distributions by token frequency and occurrence, as well as a method for validating the performance of the system.

Applying this method to conference and journal reviews written before and after the release of ChatGPT shows evidence that roughly 7-15% of sentences in ML conference reviews were substantially modified by AI beyond a simple grammar check, while there does not appear to be significant evidence of AI usage in reviews for Nature. Finally, we demonstrate several ways this method can support social analysis. First, we show that reviewers are more likely to submit generated text for last-minute reviews, and that people who submit generated text offer fewer author replies than those who submit written reviews. Second, we show that generated texts include less specific feedback or citations of other work, in comparison to written reviews. Generated reviews also are associated with lower confidence ratings. Third, we show how corpora with generated text appear to compress the linguistic variation and epistemic diversity that would be expected in unpolluted corpora. We should also note that other social concerns with ChatGPT presence in peer reviews extend beyond our scope, including the potential privacy and anonymity risks of providing unpublished work to a privately owned language model.

Conclusion. We emphasize here that we do not wish to pass a value judgement or claim that the use of AI tools for review papers is necessarily bad or good. We also do not claim (nor do we believe) that many reviewers are using ChatGPT to write entire reviews outright. Our method does not constitute direct evidence that reviewers are using ChatGPT to write reviews from scratch. For example, it is possible that a reviewer may sketch out several bullet points related to the paper and uses ChatGPT to formulate these bullet points into paragraphs. In this case, it is possible for the estimated α to be high; indeed our results in Appendix 4.6 is consistent with this mode of using LLM to substantially modify and flesh out reviews.

To enhance transparency and accountability, future work should focus on applying and extending our framework to estimate the extent of AI-generated text across various domains, including but not limited to peer review. We believe that our data and analyses can serve as a foundation for constructive discussions and further research by the community, ultimately contributing to the development of robust guidelines and best practices for the ethical use of generative AI.

Limitations. While our study focused on ChatGPT, which dominates the generative AI market with 76% of global internet traffic in the category (Van Rossum, 2024), we acknowledge that there are other diverse LLMs used for generating or rephrasing text. However, recent studies have found that ChatGPT substantially outperforms other LLMs, including Bard, in the reviewing of scientific papers or proposals (Liang et al., 2023c; Liu & Shah, 2023). We also found that our results are robust on the use of alternative LLMs such as GPT-3.5. For example, the model trained with only GPT-3.5 data provides consistent estimation results and findings, and demonstrates the ability to generalize, accurately detecting GPT-4 as well (see Supp. Table 28 and 29). However, we acknowledge that our framework’s effectiveness may vary depending on the specific LLM used, and future practitioners should select the LLM that most closely mirrors the language model likely used to generate their target corpus, reflecting actual usage patterns at the time of creation.

Our findings are primarily based on datasets from major ML conferences (ICLR, NeurIPS, CoRL, EMNLP) and Nature Family Journals spanning 15 distinct journals across different disciplines such as medicine, biology, chemistry, and environmental sciences.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do AI hiring systems affect authenticity, fairness, and candidate preferences? How can we detect and account for LLM involvement in academic writing? How can AI systems reliably guide voters without introducing political bias? What gaps exist between benchmark performance and real deployment outcomes? How do writers navigate authorship and delegation with AI?