How much peer review text shows signs of LLM modification?
Researchers analyzed AI conference reviews to estimate what fraction might have been substantially altered by large language models. Understanding this helps clarify how AI tools are entering academic peer review.
The authors estimate that between 6.5% and 16.9% of text submitted as peer reviews to ICLR 2024, NeurIPS 2023, CoRL 2023 and EMNLP 2023 "could have been substantially modified by LLMs, i.e. beyond spell-checking or minor writing updates." The number is a population share for those four venues, not a count of reviews. The authors report no comparable change in Nature family journal reviews. Where they locate user behavior is in the circumstances: the fraction is higher in reviews that report lower confidence, were submitted close to the deadline, and come from reviewers "less likely to respond to author rebuttals."
The method is what the paper calls distributional GPT quantification. It starts from a stated problem: LLM detectors have "unstable performance," so instead of classifying each review and counting, the authors fit a maximum likelihood estimate of the mixture. Reviews known to be human-written supply one reference distribution, and reviews an LLM generates from the same review instructions and papers supply the other. The vocabulary is adjectives, because a word like "commendable" spikes in recent ICLR reviews and occurs more often in generated text, with technical keywords removed. The authors report that adverbs, verbs and nouns give similar results. They also claim the approach is "more than 10 million times" cheaper computationally than state-of-the-art detectors, with estimation error reduced by factors of 3.4 in-distribution and 4.6 out-of-distribution. Those are the authors' own comparisons.
Set against the nearest notes, this is the peer-review case of the population move in How fast did LLM writing adoption actually spread?, which measures four public-facing domains. Its reason for avoiding per-document judgment matches Can people reliably spot content made by AI?, though this excerpt tests no human judges. The corpus-level compression the authors describe, where generated text narrows "linguistic variation and epistemic diversity," is the aggregate form of the lexical gap in Can human judges detect measurable differences in AI text?. The reviewing setting also meets How much does rhetorical style shift AI review scores? from the other side: there, AI reviewers' scores move with rhetoric, while here human reviewers' text is partly LLM-modified. The excerpt does not connect the two.
The excerpt does not establish that reviewers wrote reviews with ChatGPT from scratch. The authors say so directly: "Our method does not constitute direct evidence that reviewers are using ChatGPT to write reviews from scratch." A reviewer who sketches bullet points and has an LLM flesh them out would also produce a high estimate. The estimate rests on one generator family; the authors report that GPT-3.5 data generalizes to GPT-4. The discussion states the share as "roughly 7-15% of sentences," a different unit from the abstract's share of text. The confidence, deadline and rebuttal links are associations, not tested causes. The implication is narrow: for four venues in 2023 and 2024 and one generator family, a corpus-level share of 6.5% to 16.9% is consistent with substantial LLM modification, and the corpus trend is a firmer claim than any label on a single review.
Inquiring lines that read this note 12
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?- How much of ICLR 2026 peer review was already conducted by AI?
- How often do journal editors catch obvious textual problems before publication?
- Do peer review policies banning LLM use actually change reviewer behavior and decisions?
- What happens when conferences enforce bans or limits on reviewer LLM use?
- Do reviewer rules about LLM use in peer review actually get followed?
- Do conference policies banning LLM use actually reduce AI involvement in reviews?
- Do peer reviewers actually follow policies that ban or limit their LLM use?
- How effective are journal policies restricting LLM use in peer review?
- Do metareviewers and regular reviewers use LLMs differently in peer review?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How fast did LLM writing adoption actually spread?
Does LLM-assisted writing use follow a predictable adoption curve across different sectors? Understanding the speed and pattern of adoption helps explain how quickly new AI tools reshape professional communication.
sibling population estimate; this one uses explicit human and AI reference corpora on peer reviews
-
Can people reliably spot content made by AI?
This systematic review of 30 studies asks whether human judgment can distinguish AI-generated text, images, and voice from human-created content, and whether detection accuracy has improved as AI becomes more realistic.
gives the reason to avoid per-document judgment; this excerpt does not test human judges
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
same peer-review setting from the reviewer side; the two are not connected in the excerpt
-
Can human judges detect measurable differences in AI text?
Research shows LLM text differs statistically across six lexical dimensions, but human readers—even experts—cannot reliably identify which texts are AI-generated. Why does measurement succeed where human perception fails?
parallel: corpus-level compression of linguistic variation is a measurable gap at aggregate scale
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Stop Automating Peer Review Without Rigorous Evaluation
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Mapping the Increasing Use of LLMs in Scientific Papers
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
Original note title
between 6.5% and 16.9% of AI conference peer review text could be substantially LLM-modified — a corpus estimate that classifies no single review