More Versus Better, Part I

Paper · Source
Knowledge After the Web

Source: Lamar Pierce, Claudine Gartenberg, Alex Murray & Sharique Hasan, Organization Science editors' Substack · 2026-04-27

Something is up in academic research. Or in academic papers, at least.

If you are an editor or reviewer at a journal these days, you probably already know this. The manuscripts are arriving in greater volume, with a particular feel that is hard to pin down. On the surface, the papers look the same as ever, but the writing feels weightless in a way that rarely describes academic writing. The sentences connect smoothly, include citations to actual papers, and draw inferences from real (you hope) analyses. Yet you find yourself scratching your head at the meaning the words are trying to convey.

As the AI Task Force for Organization Science, we spent the past several months trying to put form to this feeling. What we find, in short, is that AI language models, in combination with strong publish-or-perish incentives, are pushing the field to produce more research rather than better research. The system is being overwhelmed by research that is, at a minimum, assisted by AI and, increasingly, substantially generated by it. This research is worse along two observable dimensions: writing quality (perhaps surprisingly) and overall quality, as measured by editorial outcomes (also perhaps surprisingly). Moreover, this research is not costless. It is imposing a burden on the volunteer labor of reviewers and editors that holds the peer review system together. Our report details what we found.

Before we get into our findings, it’s important to state up front where we (the authors of the report) stand on AI. The short version is that we are in awe of the technology. These tools have fundamentally changed our own research and teaching over the past year. Scientists have gained a real superpower, similar to moving from the slide rule to Stata 19 overnight. In that sense, we are very excited about the promise of AI for science and we are open to experimentation with it. But we are also carefully studying how it’s being used in practice and what that means for evaluating and promoting scientific research.

The problem is that AI does not appear to be changing our field for the better so far. Instead, Organization Science, like many other journals, is overwhelmed by AI-generated research that is stressing its peer review system to the breaking point. First, let’s start with raw numbers. Submission volume at Organization Science has risen by 42% since the launch of ChatGPT in November 2022, compared to the prior two-year window. To put that in perspective, the COVID-19 pandemic, which sent many academics home with canceled conferences and more time to write, produced a 20% bump. The post-AI surge is on top of this base, and continues to increase.

The increase in submissions is almost entirely due to manuscripts with substantial AI-generated text. Submissions with little or no detectable AI (below 15% on Pangram, a widely used AI detection tool) have actually declined since late 2022. The gap between that decline and the 42% aggregate increase is filled by manuscripts in the higher AI categories: moderate AI collaboration (15-30%), substantial AI generation (30-70%), and what we consider mostly AI-written work (70%+). By early 2026, the majority of manuscripts submitted to Organization Science contain detectable AI-generated writing. The fastest-growing category is also the most concerning: manuscripts where 70% or more of the text appears to be AI-generated.

To be clear about what we are measuring: we used Pangram to score AI prevalence in writing on a continuous scale from 0 to 100% for each submission’s abstract. We validated this approach against full manuscripts and found that high AI scores in the abstract reliably signal heavy AI use throughout the paper, across the introduction, theory, methods, results, discussion, and conclusion sections. Based on our data, an abstract that reads like it was written by a language model is a decent, though not perfect, indicator that the rest of the manuscript was too.

A natural response to these numbers is: so what? If AI helps researchers write better papers faster, more submissions could be a good thing. The field has long worried about barriers to entry, about non-native English speakers facing disadvantages, and about the sheer drudgery of academic prose. If AI helps with these, the growth we are witnessing might carry good work in with it and more equitable opportunities. But that’s not what seems to be happening, at least from an exposition standpoint.

We calculated the average readability of submission abstracts over more than a decade. Writing quality was stable from the beginning of this period through COVID. Then, right at the launch of ChatGPT, readability scores began to drop. By January 2026, the average abstract’s Flesch Reading Ease score is 1.28 standard deviations lower than it was in January 2021. Submissions have become far harder to read.

This is counterintuitive. Most people assume that AI produces cleaner, more polished text. And in some narrow dimensions, it does: AI-assisted writing tends to be less hedging, less passive, and more specific (for instance, it includes more numbers). But on the measures that capture whether a reader can actually parse and absorb the prose, AI writing is worse. AI-generated academic writing uses longer words, more complex sentence structures, more jargon, and more nominalizations (i.e., those nouns that we never learned until we entered academia such as “conceptualization,” “problematization,” “contextualization,” and so forth).

The result is text that is superficially clear but substantively impenetrable.

To test this ourselves, Claudine took Lamar’s most recent publication (while Lamar was distracted by feeding Butterburger), stripped it of the abstract, title, and introduction, and asked ChatGPT 5.4 to write a title and abstract, aimed at a “top” journal with “expert” editors and reviewers as the audience. Here’s Lamar’s original followed by ChatGPT’s version.

“Behavioral economists have proposed that incentive contracts result in higher productivity when bonuses are “loss framed”—prepaid then clawed back if targets are unmet. We test this claim by randomizing the pre- or post-payment of sales bonuses at 294 car dealerships. Although somewhat statistically imprecise, our analysis provides strong indications that the random assignment of loss framing had quantitatively important negative effects. We document that the negative effects of loss framing can arise due to an increase in incentives for “gaming” behaviors. Based on these claims, we reassess the common wisdom regarding the desirability of loss framing.”

“Loss framing is widely assumed to enhance effort and performance, yet field evidence remains limited and theoretically incomplete. We report results from a large-scale randomized controlled trial involving automobile dealerships, in which a standard bonus contract was experimentally reframed from gains to losses via upfront payments with clawbacks. Contrary to prevailing expectations, loss framing weakly reduces overall performance, with substantial heterogeneity across task dimensions. Specifically, we find no effect on high-stakes sales categories but significant declines in lower-stakes categories, driven by shifts in the distribution of outcomes rather than uniform changes in effort. To explain these patterns, we develop a model of reference-dependent agents engaged in multidimensional tasks. The model shows that loss framing amplifies incentives to minimize loss exposure, which can induce inefficient reallocation of effort across tasks or toward lower-variance strategies. Consistent with this mechanism, distributional analyses reveal both increased target attainment in focal domains and offsetting declines elsewhere. Our findings challenge the generalizability of prior results on loss framing and highlight a critical boundary condition: when performance involves trade-offs across activities, loss-framed incentives may distort rather than enhance productivity.”

The ChatGPT version makes sense, is grammatically correct, and yet is a slog to get through and much harder to grasp than the original.

Now, extend these effects to the full manuscript. When AI scores reach above 30% and particularly above 70%, it means that the manuscript is similarly a slog to understand. It also implies that, during the research process itself, authors substantially delegate writing, and the thinking that goes along with it, to the models. The aha moments that come from the writing process are now gone.

If the new submissions consisted of strong papers, it would be a welcome problem. More good ideas competing for journal space is only a good thing. But the editorial process tells us that these AI-heavy submissions are, on the whole, weaker manuscripts.

Among manuscripts with 70%+ AI scores, nearly 70% are desk-rejected, meaning an editor determined they should not be sent out for review.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do hallucinated citations emerge in AI scholarly output? Can AI systems perform peer review as effectively as humans? Does disclosing AI authorship change how audiences evaluate the writing? How do educators verify student capability when AI can produce indistinguishable work? How do AI hiring systems affect authenticity, fairness, and candidate preferences? Can readers reliably distinguish AI-written text from human writing? How do writers navigate authorship and delegation with AI? What are the real-world consequences of AI citation hallucinations? How reliably can humans and AI detectors identify machine-generated text? Are AI-generated articles systematically disadvantaged in search ranking and user engagement? Why do confident AI outputs mislead human trust calibration?