How much scientific writing has LLMs actually modified?
Researchers estimated the fraction of LLM-modified content across nearly one million scientific papers from 2020 to 2024. Understanding these population-level trends matters for gauging AI's real impact on scientific publishing.
The authors estimate that LLM-modified content rose steadily across 950,965 papers posted to arXiv, bioRxiv and 15 Nature portfolio journals between January 2020 and February 2024, with the largest and fastest growth in Computer Science. The abstract gives the Computer Science figure as "up to 17.5%" and the least modification, in Mathematics and the Nature portfolio, as "up to 6.3%". The corpus is 773,147 arXiv papers, 161,280 bioRxiv papers and 16,538 Nature portfolio papers. The discussion describes "a sharp increase in the estimated fraction of LLM-modified content" beginning about five months after ChatGPT's release. These are the authors' own estimates from their own method, not counts of tagged papers.
The method is a distributional framework adapted from Liang et al. (2024). It models how often each token appears in human-written and LLM-modified text and recovers the LLM-modified fraction across a corpus, so no individual paper is classified. The authors make two changes. Training text is generated counterfactually: an LLM summarizes a human-written paragraph into an outline, then writes a full paragraph from it, which the excerpt says "simulates how scientists may be using LLMs in the writing process". They also estimate from the full vocabulary rather than adjectives alone, since adjectives, adverbs and verbs all performed well in their validation.
This is the scientific-publishing instance of the population-level design behind How fast did LLM writing adoption actually spread?, which applies the same logic to consumer complaints, press releases, job postings and UN releases. The two series differ in reach. The four-domain estimate shows a surge and then a plateau through September 2024; this excerpt's series ends in February 2024 and describes a steady rise, so it cannot test the question raised in Is the 2024 LLM writing plateau real saturation or measurement artifact?. The design also explains why measurement works where detection fails. The introduction says less obvious cases are "nearly impossible to detect at the individual level", and Can human judges detect measurable differences in AI text? shows the same split between aggregate measurability and human perception. The reader study in Can readers tell LLM abstracts from human ones? asks judges about abstracts; this excerpt measures the same kind of text at corpus scale and asks no one to judge.
The excerpt does not establish why growth differs by field. Its two explanations, Computer Science researchers' "familiarity with and access to large language models" and "the pressure to publish quickly", are offered but not tested. The associations with preprint posting, crowded research areas and shorter papers are correlational, and the excerpt does not separate them. The estimator's validation is inherited from Liang et al. (2024) and not reproduced here. The limitations section's concern about "the security and independence of scientific practice" is raised, not measured. The implication is narrower than the headline: LLM modification of abstracts and introductions rose across this corpus from about five months after ChatGPT's release, most in Computer Science, and the causes of that gap are open.
Inquiring lines that read this note 2
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can we detect and account for LLM involvement in academic writing?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How fast did LLM writing adoption actually spread?
Does LLM-assisted writing use follow a predictable adoption curve across different sectors? Understanding the speed and pattern of adoption helps explain how quickly new AI tools reshape professional communication.
same population-level design in four non-scientific domains; this series is scientific and ends before their plateau.
-
Can human judges detect measurable differences in AI text?
Research shows LLM text differs statistically across six lexical dimensions, but human readers—even experts—cannot reliably identify which texts are AI-generated. Why does measurement succeed where human perception fails?
aggregate measurability without human detectability, the gap the population design relies on.
-
Can readers tell LLM abstracts from human ones?
Do readers with ML expertise reliably distinguish human-written, LLM-generated, and LLM-edited research abstracts? Understanding this matters for evaluating whether readers can serve as effective gatekeepers against LLM content.
same text type, abstracts, judged by readers there and measured at corpus scale here.
-
Is the 2024 LLM writing plateau real saturation or measurement artifact?
The adoption curve for LLM-assisted writing flattened in 2024, but the cause remains unclear: either genuine saturation or models becoming too subtle to detect. Resolving this matters for understanding actual usage trends versus measurement limitations.
this series stops in February 2024 with a steady rise, so it cannot resolve the plateau question.
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Mapping the Increasing Use of LLMs in Scientific Papers
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- LLM-Generated or Human-Written? Comparing Review and Non-Review Papers on ArXiv
- Scientific production in the era of Large Language Models
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Has the Creativity of Large-Language Models peaked? —an analysis of inter- and intra-LLM variability —
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026
Original note title
estimated LLM-modified content rose steadily across 950,965 scientific papers from 2020 to 2024, fastest in computer science at up to 17.5%