Screening, sorting, and the feedback cycles that imperil peer review
Scholarly journals rely on peer review to identify the science most worthy of publication. Yet finding willing and qualified reviewers to evaluate manuscripts has become an increasingly challenging task, possibly even threatening the long-term viability of peer review as an institution. What can or should be done to salvage it? Here, we develop mathematical models to reveal the intricate interactions among incentives faced by authors, reviewers, and readers in their endeavors to identify the best science. Two facets are particularly salient. First, peer review partially reveals authors’ private sense of their work’s quality through their decisions of where to send their manuscripts. Second, journals’ reliance on traditionally unpaid and largely unrewarded review labor deprives them of a standard market mechanism—wages—to recruit additional reviewers when review labor is in short supply. We highlight a resulting feedback loop that threatens to overwhelm the peer review system: (1) an increase in submissions overtaxes the pool of suitable peer reviewers; (2) the accuracy of review drops because journals either must either solicit assistance from less qualified reviewers or ask current reviewers to do more; (3) as review accuracy drops, submissions further increase as more authors try their luck at venues that might otherwise be a stretch. We illustrate how this cycle is propelled by the increasing emphasis on high-impact publications, the proliferation of journals, and competition among these journals for peer reviews. Finally, we suggest interventions that could slow or even reverse this cycle of peer-review meltdown.
Introduction. I. INTRODUCTION When we think about the institution of peer review in science, we often envision reviewers as acting akin to shell collectors, sorting through the fragments of shells on a tropical beach in search of whole specimens worth taking home. Indeed, this is part of what peer review does. But papers don’t simply appear on Nature’s editorial desk the way that shells wash up on an expanse of white sand. Authors choose to put them there—and they make those choices in anticipation of the evaluation to follow. Given the stringent review process and the wasted time and effort involved in submitting a paper that is eventually rejected, authors screen their own work, targeting appropriate journals rather than sending everything to the most prestigious outlets. Thus peer review also induces authors to reveal, through their submission decisions, their own private information about the quality of their work [1–3]. However, this service depends on a massive supply of free labor, namely the unpaid and largely uncredited efforts of peer reviewers [4]. But the pool of peer-review labor has become stretched thin, and the problem seems to be worsening. Editors report increasing difficulty in finding reviewers for the manuscripts that they handle. Bibliometric studies in a range of fields support their assertions: reviewers are more likely to decline review invitations and thus the average number of solicitations per acceptance has increased [[5–9], though see [10] for an exception]. We appear to be in the midst of a “peer-review meltdown” in which the peer-review system is becoming woefully overtaxed by the volume of manuscript submissions [11–16].
As peer review teeters, scientists have begun experimenting with new ideas to reduce the review load or increase review supply. For example, venues such as Publons and Elsevier’s Reviewer Recognition Platform attempt to make reviewing more prestigious by awarding accolades to top reviewers [17, 18]. Brokerage services tried charging authors to secure peer reviews that could be forwarded to prospective publication outlets [19, 20]. These foundered, but a new generation of journal-independent review initiatives such as Review Commons and Peer Community In have emerged in their stead. Some computer science conferences such as NeurIPS [21] keep review loads down by disallowing revision and re-review. Elsewhere, some journals offer portable peer review, where initial reviews follow a manuscript if it is resubmitted to other venues. Others have suggested tying the opportunity to submit a paper as an author to one’s contributions as a reviewer [22], and some journals have experimented with cash payments for reviews [23–26]. Yet others have studied whether machine review using AI [27] and large language models [28] can complement peer review. While debate about the propriety of machine review remains unsettled [29, 30], some reviewers are using LLMs for assistance even when journal or conference guidelines forbid doing so [31–33]. Some critics have even proposed eliminating prepublication peer-review altogether [34]. The breadth of these endeavors testifies to scientists’ eagerness to place scientific publishing on more stable footing. Yet all these initiatives are hampered by the fact that a rigorous theory of the structure and function of peer review has yet to coalesce [35]. This article aims to begin to fill that gap. While the literature is dotted with mathematical models of peer review, many of these efforts use detailed, agent-based simulation models that embrace the richness of the scientific ecosystem (e.g., [36–40]). In this paper, we take a different approach by developing low-dimensional models that isolate how peer review shifts the burden of identifying the best science among among authors, reviewers, and readers of the scientific literature. (Zhang et al. [41] have recently presented a model of computer-science conferences that also examines how self-screening by authors affects the peer-review burden.) These models reveal a set of hidden pressures on peer review, and bringing them to light helps us to understand the tensions that threaten the entire institution as science changes. In particular, these models highlight a pernicious feedback loop: increasing submissions to top journals exhausts the pool of suitable peer reviewers, resulting in lower quality peer reviews that encourage yet more authors to take a chance on submitting their paper to a prestigious venue (Fig. 1). We also show how other systemic factors—increasing emphasis on high-impact publications, the proliferation of journals, and competition among these journals for unpaid peer-review labor—propel this cycle. Finally, we consider possible solutions that could interrupt the meltdown of peer review, or even reverse it.
Method. II. MODEL AND WELFARE A. Adda-Ottaviani model Our model builds from the base model in Adda & Ottaviani (2024; henceforth AO) [42]. AO use their model to study grant competitions, but their model adapts naturally to scientific publication. Our analysis differs substantially from AO, as suits the different aims of the papers. The appendix provides mathematical proofs of several of the key claims and some additional results. Consider a scientific community served by two journals: an elite journal that seeks to publish top manuscripts and a mega-journal that publishes everything else. We consider only two journals for the sake of simplicity, although our model applies in any setting FIG. 1. The peer-review meltdown cycle. Author screening and journal sorting interact in a feedback loop which inaccurate sorting loosens author screening [42] and looser screening makes sorting less accurate by depleting the pool of available review labor. Forces that exacerbate this feedback loop include (clockwise from top): Greater rewards to publishing lead more authors to submit their paper to top journals; a proliferation of journals gives authors more opportunities to obtain fresh reviews of already rejected manuscripts; journals’ reliance on a shared pool of review labor compels journals to underuse desk rejection and overexploit the review pool; and noisier review reduces what authors can learn from having their papers rejected. with several vertically differentiated journals. Suppose that the elite journal has the capacity to publish a proportion k ∈(0, 1) of the manuscripts that the community produces. We assume that the journal’s capacity k is determined exogeneously, perhaps by limits imposed by the publisher or by constraints on readers’ attention [43]. Henceforth, we focus on the behavior of the elite journal and place the mega-journal in the background. Thus we refer to the elite journal as simply “the journal”; we say that authors who publish their manuscripts in the elite journal are “published”, and so on. Suppose that this community contains a unit mass of authors and that each author is endowed with a manuscript with quality θ. For mathematical convenience, assume that θ has a standard Gaussian distribution across authors, θ ∼N(0, 1). Authors have their own sense of whether their work is any good. We instantiate this by assuming that an author with a manuscript of quality θ obtains a private sense of their manuscript’s quality X that is drawn from a normal distribution with mean θ and variance σ2 X. Across authors, X is marginally distributed as N(0, 1 + σ2 X). We refer to the quantile q = FX(x) as the author’s type (here and throughout, F denotes a cumulative distribution function, or cdf), and identify authors with their type, e.g. “author q”. Authors can submit their manuscript to the journal, or not. Authors who submit their manuscript pay a disutility cost c > 0 [3], which includes the opportunity cost of foregoing immediate publication in the mega-journal. Authors whose manuscripts are published receive kudos, prestige, professional rewards, etc. with value v > c. More precisely, v gives the additional reward that the author receives from publishing in the elite journal right away relative to the time-discounted reward of publishing in the mega-journal later. We ignore any other actions that authors may take (e.g., cover letters) that could communicate information about their type to the journal. Because the journal cannot directly observe manuscript quality θ, it solicits reviews in the usual way. For now, assume that journals send out every manuscript they receive for review; desk rejection will be considered later. Let Y be the review score for a submitted manuscript, and assume that for a manuscript of quality θ, Y is drawn from a Gaussian distribution with mean θ and variance σ2 Y . Higher review scores Y provide evidence of higher article quality θ, and thus the journal rationally publishes those papers whose review score exceeds some acceptance threshold y. Two conditions determine the model equilibrium. First, authors submit their paper if and only if their payoff from doing so is positive.
Discussion. IV. DISCUSSION Some of the drivers of the peer-review crisis are straightforward and need little elaboration. The number of papers published has been growing about 5% annually since 1952 [53, 54], a rate that exceeds the concomitant growth in the number of university faculty [55]. Of those faculty, a growing proportion are employed in short-term contingent positions and may be less able to devote time to unpaid and unrecognized review service. Rapid growth of scientific productivity in non-Western countries has outpaced the fraction of review invitations going to authors located in these regions [7], presumably because of mismatches in the composition of editorial boards. These trends have caused peer-review effort to become imbalanced, with a comparably small fraction of researchers providing most of the review labor [56, 57]. Meanwhile, both the intrinsic rewards of reviewing and the cost of declining review invitations have shifted. Preprint servers and their ilk have eliminated what used to be one of reviewing’s major perks, namely the opportunity to obtain an early look at important unpublished work. Scientific communities are larger and looser knit, making the professional networks that bind editors and reviewers more diffuse. The universal reliance on e-mail for communication between journals and reviewers has made editors’ invitations weaker signals of genuine interest in a reviewer’s opinion compared to the manila envelopes and hand-written notes of yesteryear, while evolving social norms reduce the psychic cost to reviewers of ignoring e-mailed invitations altogether. Other forces driving the peer-review crisis are subtle and more complex. We have argued here that the peerreview crisis is partially driven by a pernicious feedback loop in which a growing burden on the reviewer community leads to less accurate reviewing, which in turn compels authors to take more chances with regard to where they submit their work, which exacerbates the burden on the reviewer community further, and so on. Moreover, the decentralized nature of scientific publishing and the need for journals to compete for authors, readers, and reviewers short-circuits many of the most obvious paths towards halting or reversing this cycle. While we have presented our model in the context of elite journals, that focus is merely a simplification to streamline the exposition. Similar forces buffet journals across the scientific spectrum. Like all models, our model makes simplifications that provide scope for future work. Perhaps the most substantively, we treat manuscript quality as exogenous. Of course, authors are not endowed with fully formed manuscripts, but instead they decide how much effort to invest to developing manuscripts based on their anticipation of the journal’s scrutiny. For example, the rise of revise-and-resubmit as an author’s likely best outcome may perversely encourage authors to submit manuscripts that are less than polished, both because revision seems inevitable and because journals can no longer expect polished submissions if unpolished submissions become the norm [58]. A model of journal and author behavior that endogenizes manuscript quality would provide more satisfying insight on this count. If contemporary science has reached a point in which the demand for review labor outstrips the available supply, what (besides increasing desk rejection) might journals do to reconcile the two? Journals might consider tightening screening by increasing the cost of preparing a manuscript, c, thus discouraging more authors from submitting their manuscripts in the first place [2].
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems perform peer review as effectively as humans?- How often do false positives from detection tools actually occur in peer review?
- Why does publish-or-perish incentivize quantity over quality in research?
- Why do peer reviewers favor novel ideas that later fail in execution?
- Do citation counts better capture scientific quality than publication venue tiers?
- Why do individual peer reviewers show such low agreement on research merit?
- Why do authors submit manuscripts to venues beyond their reach?
- How fast is scientific publishing growing relative to reviewer capacity?
- Do shortened peer review timelines correlate with lower quality publications?
- What makes disruptive scientific work harder to publish and recognize?
- What effects do preprint servers have on scientific consensus formation?
- Can institutional statements alone correct misconceptions from unreviewed papers?
- Can automated systems scale peer review faster than human moderators?
- Should citation counts serve as the primary measure of research impact?
- Do preprint servers have tools to detect hidden text in submitted manuscripts?
- What role do conference organizers play in accepting problematic articles?
- Does the form of a paper still matter if the process behind it changes?