Scientific production in the era of Large Language Models
Large Language Models (LLMs) are rapidly reshaping scientific research. We analyze these changes in multiple, large-scale datasets with 2.1M preprints, 28K peer review reports, and 246M online accesses to scientific documents. We find: 1) scientists adopting LLMs to draft manuscripts demonstrate a large increase in paper production, ranging from 23.7-89.3% depending on scientific field and author background, 2) LLM use has reversed the relationship between writing complexity and paper quality, leading to an influx of manuscripts that are linguistically complex but substantively underwhelming, and 3) LLM adopters access and cite more diverse prior work, including books and younger, less-cited documents. These findings highlight a stunning shift in scientific production that will likely require a change in how journals, funding agencies, and tenure committees evaluate scientific works.
Introduction. The scientific enterprise is intimately connected with technological innovation. The microscope (1), advances in computing (2, 3), and next-generation sequencers (4), for example, shifted the frontier of research. Today, the fast adoption of generative artificial intelligence (Gen AI) across all academic disciplines (5-8) is recasting scientific production. Despite growing excitement (and concern) about Gen AI’s role in research, empirical evidence remains fragmented, and systematic understanding of the impact of Large Language Models (LLMs) across scientific domains is limited.
Researchers have demonstrated the value of AI in many specific scientific contexts (9-11), such as protein structure prediction (12) and materials discovery (13). Recent advancements in LLMs have expanded its use across a wide range of tasks in natural (14-16) and social sciences (17-20). This work highlights the incredible potential of LLMs across specific scientific undertakings, raising an open question: What is the macro level impact of LLMs on the scientific enterprise?
Method. To address this question, we collected large-scale data from three preprint repositories (Jan 2018 to June 2024, see SM S1.1-1.3 for details): (1) arXiv (1.2M preprints), which includes mathematics, physics, computer science, electrical engineering, quantitative biology, statistics, and economics, (2) bioRxiv (221K preprints), which spans a wide range of subfields in biology and the life sciences, and (3) Social Science Research Network (SSRN, 676K preprints), a working paper repository hosting manuscripts in the social sciences, law and the humanities. Each of the three datasets represents the largest within its domain. Collectively, they offer an unprecedented empirical basis to examine some of the impacts of LLMs on scientific productivity practices across many scientific fields.
To identify the use of LLMs in the creation of scientific manuscripts, we developed a text-based AI detection algorithm (5), which we applied to all abstracts in our data. We used abstracts from papers submitted prior to 2023 – before the ChatGPT era – to estimate the token distribution of human-written text, then prompted OpenAI’s GPT-3.5turbo0125 model to rewrite these abstracts to generate the token (word) distribution of LLM-written text. We compared these distributions to identify probable LLM-assisted abstracts written after the release of ChatGPT 3.5. Further details on model training, validation, potential limitations, and alternative methods of LLM detection are provided in SM S2.1, S4, and S5.
Discussion. A productivity jump may stem from the use of Gen AI across multiple research tasks, including idea generation, literature discovery, coding, data collection or analysis. But to date, LLMs likely have had the largest impact in writing (24, 25). To create distinctive scientific works, researchers must present compelling written arguments; link a manuscript’s arguments, methods, and results to prior literature; detail and contextualize the most important findings; and articulate what can be learned from the text. These complex writing tasks are time consuming, particularly for researchers communicating in a non-native language. We therefore ask: Does the productivity impact of LLM adoption vary across authors’ native language proficiencies? Since most high-impact research is published in English-language journals and proceedings, native speakers have had a substantial advantage in scientific communication (26-28). LLMs can mitigate disparities in English fluency, which should asymmetrically reduce the cost of writing across scientists’ linguistic backgrounds (29).
We conclude that even the use of previous-generation LLMs 3⁄4 those available to scholars at the time the manuscripts in our data were drafted 3⁄4 are associated with productivity gains, particularly for researchers facing higher costs of writing. These findings concur with work showing that LLMs mitigate the impact of skill disparities, in this case by reducing the cost of writing in a second language (30). Given considerable advances in the writing ability of currentgeneration LLMs and more widespread availability of these systems, the productivity effects we estimate are likely substantial enough to imply a shift in the market share of scientific production toward scholars in non-native English-speaking geographies.
LLMs are likely to reshape science production beyond productivity effects. High-quality writing is often construed as a signal of scientific merit. Papers with clear but complex language are perceived to be stronger and are cited more frequently (31). Because novel scientific advances are the product of years of knowledge refinement, the ability to precisely articulate scientific discoveries is a (very imperfect) proxy for the care of a scientific team’s work. The fact that LLMs can almost effortlessly produce polished, professional text describing any scientific topic raises an important question: Does LLM use reveal or conceal the quality of the underlying research?
The sharp contrast in quality assessments across the distribution of language complexity in the two groups 3⁄4 human-written and LLM-assisted manuscripts 3⁄4 confirms that complex LLMgenerated language often disguises weak scientific contributions (35). These findings demonstrate the rapid erosion of a traditional heuristic. For LLM-assisted manuscripts, the positive correlation between linguistic complexity and scientific merit not only disappears; it inverts. As the effort required to produce polished prose declines, so too does its utility as a signal of an author’s command of a topic (36). This creates a risk for the scientific enterprise, as a deluge of superficially convincing but scientifically underwhelming research could saturate the literature. If this occurs, it will bury some important ideas and force the community to waste valuable time separating genuine insights from a morass of unimportant work.
For peer reviewers and journal editors, this represents a significant issue. As a shortcut to (imperfectly) screen scientific research, writing characteristics are fast becoming uninformative signals just as the quantity of scientific communication surges.
Conclusion. Our findings show that LLMs have begun to reshape scientific production. Use of LLMs accelerates manuscript output, reduces barriers for non-native English speakers, and diversifies the discovery of prior literatures. These changes may democratize scientific production. However, traditional signals of scientific quality such as language complexity are becoming unreliable indicators of merit just as we experience an upswing in the quantity of scientific work. These changes portend an evolving research landscape in which the value of English fluency will recede, but the significance of robust quality assessment frameworks and deep methodological scrutiny is paramount. As AI systems advance, they will challenge our fundamental assumptions about research quality, scholarly communication, and the nature of intellectual labor. An urgent agenda for science policy makers will be to evolve our scientific institutions to accommodate the rapidly evolving scientific production process.
Limitations. This study explores the impact of LLMs on scientific production, but our findings are subject to several limitations that offer avenues for future research.
First, we do not provide causal identification. A central difficulty in studying LLMs “in the wild” is the impossibility of perfectly measuring their use. Our AI detection method is imperfect and susceptible to several challenges (SM S5.1-5.4): it relies on abstracts rather than full text (SM S5.5); it cannot definitively identify which specific co-author on a team used an LLM (SM S5.6); and it almost certainly fails to detect use by authors who heavily edit LLM-assisted text. Furthermore, the non-random adoption of Gen AI tools creates the potential for self-selection bias, and our focus on published papers means the “adoption time” may be endogenous to productivity. Our supplement contains many additional analyses to evaluate the scope of these issues, and while our results appear robust, it is important for future work to continue to identify methodological strategies to address these challenges.
Second, our findings represent a snapshot of a rapidly evolving technology. Our analysis is based on data generated prior to the arrival of more advanced reasoning models and deep research capabilities.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How can we detect and account for LLM involvement in academic writing?- Why do computer science papers show more LLM modification than other fields?
- Did LLM use in scientific writing plateau after the initial surge?
- What concerns does widespread LLM use raise for scientific independence?
- Does the form of a paper still matter if the process behind it changes?
- Can statistical filtering plus narrative generation fool academic peer review?
- Why does peer review fail on unrepeatable AI-generated outputs?
- Why does automated evaluation consistently overestimate research quality?
- How do LLMs generate false citations that sound like real scholarship?
- Can citation practices work when AI cannot produce traceable sources?
- How does treating synthetic data as empirical evidence contaminate statistical inference?
- How do retrieval failures enable generation of fabricated scholarly constructs?
- Can verification mechanisms prevent AI agents from inventing false citations?
- Can we verify fabricated text without redesigning the generation process?
- Can fabrication of content serve productive purposes in prediction?
- What makes counterfeiting social warrant different from counterfeiting factual claims?
- Can discourse-level analysis detect deception better than individual word choices alone?
- How do verification labels themselves become part of the misinformation problem?