How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review

Paper · arXiv 2510.02143 · Published October 2, 2025
Domain Specialization in LLMs

Peer review in academic research aims not only to ensure factual correctness but also to identify work of high scientific potential that can shape future research directions. This task is especially critical in fast-moving fields such as artificial intelligence (AI), yet it has become increasingly difficult given the rapid growth of submissions. In this paper, we investigate an underexplored measure for identifying high-impact research: authors’ own rankings of their multiple submissions to the same AI conference. Grounded in game-theoretic reasoning, we hypothesize that self-rankings are informative because authors possess unique understanding of their work’s conceptual depth and long-term promise. To test this hypothesis, we conducted a large-scale experiment at a leading AI conference, where 1,342 researchers self-ranked their 2,592 submissions by perceived quality. Tracking outcomes over more than a year, we found that papers ranked highest by their authors received twice as many citations as their lowest-ranked counterparts; self-rankings were especially effective at identifying highly cited papers (those with over 150 citations). Moreover, we showed that self-rankings outperformed peer review scores in predicting future citation counts. Our results remained robust after accounting for confounders such as preprint posting time and self-citations. Together, these findings demonstrate that authors’ self-rankings provide a reliable and valuable complement to peer review for identifying and elevating high-impact research in AI.

Introduction. A central objective of peer review is to identify and promote research of high impact that will shape the future of a field. Yet, this process is under unprecedented strain, particularly in rapidly expanding areas such as artificial intelligence (AI), which face an explosion in paper submissions (Lipton and Steinhardt, 2019; Sculley et al., 2019; PaperCopilot, 2025). For the two largest AI conferences, submissions to the International Conference on Machine Learning (ICML) rose from 1,676 in 2017 to 12,107 in 2025, while those to the Conference on Neural Information Processing Systems (NeurIPS) increased from 3,240 in 2017 to 21,575 in 2025. This exponential growth has placed a heavy burden on peer review at AI conferences, leading to reliance on graduate and undergraduate students as reviewers (Stelmakh et al., 2021), many without prior publications at these venues, and increasingly, the use of language models for review (Liang et al., 2024). This raises serious concerns about the quality of the review. A striking example is the NeurIPS 2021 experiment, which revealed that approximately half of the accepted papers would have been rejected under a second independent review (Cortes and Lawrence, 2021; Beygelzimer et al., 2023). This inconsistency highlights the challenge of identifying research with lasting impact amidst the overwhelming volume of incremental work. A consequence of the decline of peer review that we cannot preclude is that the advancement of AI may follow a less-than-optimal path. While the supply of qualified reviewers is strained, far less research has focused on eliciting expert assessment from the authors themselves (Aziz et al., 2019; Mattei et al., 2020; Srinivasan and Morgenstern, 2021; Su, 2021; Rastogi et al., 2024). Authors possess an unparalleled understanding of their own work, encompassing its theoretical foundations, limitations, and the subtle details fundamental to its long-term scientific impact. This deep insight stands in contrast to the focus of many overburdened reviewers who, under tight deadlines, may prioritize quantifiable gains over a submission’s broader scientific implications. This is particularly true in AI research, where many reviewers are junior researchers who often focus on marginal improvements in predictive accuracy (such as a benchmark test accuracy increasing from 91% to 92%) and may lack the experience to evaluate a paper’s long-term significance (Shah et al., 2018). To empirically test the hypothesis that author insight predicts scientific impact, a primary obstacle is soliciting candid self-assessments, as authors may be inclined to inflate the quality of their work. One approach is to elicit a self-ranking of submissions from the same author, which is practical as it is common for the same author to submit multiple papers to the same conference in AI research. This design elicits comparative rather than absolute judgments, so authors cannot simply claim all their submissions are of the highest quality. Under certain conditions, this comparative design incentivizes authors to truthfully report their true rankings (Su, 2021, 2025; Yan et al., 2025). To empirically quantify the predictive power of authors’ self-rankings, we designed and implemented a large-scale experiment at ICML in 2023, with the approval of the conference organizers. We asked authors who had submitted multiple papers to rank them according to their perceived quality and significance. The experiment yielded a rich dataset of self-rankings from 1,342 AI researchers, encompassing 2,592 submissions to ICML 2023, along with their official review scores and acceptance decisions. Our analysis of the ICML 2023 data reveals that authors’ self-rankings are a powerful predictor of a paper’s future impact, measured by citations accumulated over 16 months. Citation count is a widely used, albeit imperfect, metric for the scientific impact of research (Aksnes et al., 2019; Huang et al., 2022). Submissions ranked highest by their authors received, on average, twice as many citations as those they ranked lowest, a trend that held for both accepted and rejected submissions, suggesting that self-rankings can not only denoise review scores but also capture a distinct dimension of impact (Su et al., 2025). The predictive power was particularly striking for identifying high-impact work: of the 22 papers in our dataset that garnered over 150 citations, 17 were ranked highest by their authors. For comparison, we demonstrate that these self-rankings are a more accurate predictor of future citations than the review scores. These findings are highly statistically significant and remained robust after controlling for potential confounding factors such as early dissemination on preprint servers and authors’ self-citations.

Method. 2 Data collection Our experiment was conducted in two phases. In Phase One, we carried out a survey-based experiment on OpenReview, which hosted the ICML 2023 peer review process for 6,538 submissions from 18,535 authors. The experiment itself was implemented with OpenRank.cc, a platform we developed for this purpose. On January 26, 2023, immediately after the submission deadline, an official email was distributed via OpenReview to all ICML authors requesting information about their submissions. Authors with multiple submissions were asked to rank their papers based on perceived quality, along with answering several questionnaire items. In total, 5,634 authors completed the survey (response rate: 30.4%), of which 1,342 authors with multiple submissions provided rankings. Altogether, 2,592 submissions were ranked by at least one author, accounting for 39.6% of all submissions. It is common for AI researchers to submit multiple papers to the same conference. Indeed, at ICML 2023, 4,505 of the 18,535 authors had more than one submission, and 5,035 of the total 6,538 submissions had at least one author having more than one submission. A paper submitted to an AI conference is evaluated by multiple reviewers. Before the rebuttal phase, every reviewer assigns a pre-rebuttal review score—at ICML 2023, on a 1–10 scale. During the rebuttal period, authors may respond to reviewers’ comments and clarify points without making substantial changes to the original submission. After considering the rebuttal, reviewers may revise their scores; these revised scores are referred to as post-rebuttal scores. We obtained the pre- and post-rebuttal scores, as well as the final decisions for all submissions, from OpenReview. In Phase Two of the experiment, we used Semantic Scholar to collect citation data for each ranked submission.1 For each submission, we retrieved all papers that cited the submission between July 23, 2023 and November 22, 2024. For each citing paper, we extracted the title, author list, and publication date. A key challenge was accounting for multiple submission versions, which often involved revised titles or author lists before or after ICML 2023. Additionally, some titles or author names include special characters, such as “Schr ̈odinger,” which introduced minor mismatches during automated title matching. To ensure accuracy, we retained only submissions with an exact match in both title and author list on Semantic Scholar (see Section A for details). After filtering, we obtained a final dataset consisting of 797 authors with multiple ranked submissions, each with valid citation data, involving a total of 1,527 unique submissions. All subsequent citation analyses are based on this filtered set. Across these submissions, the average number of citations per paper during the study window was 14.73. Half of these submissions received at least 4 citations, and one-quarter received at least 11 citations. The most-cited submission received 979 citations.2

Discussion. Recognizing that peer review is strained by a limited reviewer pool in the search for high-impact research, our study reveals that a powerful and overlooked source of insight is the authors themselves. We demonstrated through a large-scale experiment at one of the largest AI conferences that authors’ comparative assessments of their own papers are a remarkably strong predictor of future scientific impact. Submissions that authors ranked highest garnered more than double the citations of those they ranked lowest; this signal was especially pronounced in the tail, identifying a large majority of the most highly cited papers. Crucially, self-rankings are more predictive of future citation counts than reviewers’ scores. These findings are statistically robust and persisted after controlling for confounding factors like preprint posting dates and self-citations. Citation counts are usually viewed as a strong indicator of a paper’s impact in AI. At AI conferences, papers typically introduce fresh ideas—either novel insights into existing models or entirely new methods—and there isn’t a single “ground truth” that immediately proves which approach is best; instead, the community tests, debates, and builds on them. In that context, a paper’s citation count serves as a practical signal of scientific impact: when many researchers cite a paper, it means the idea has been noticed, scrutinized, and found useful enough to inform further work, making it a reasonable proxy for scientific influence. Beyond academia, AI papers can also shape products and markets. Popular, highly cited ideas tend to attract developer mindshare, funding, and tooling—so the “hotter” the idea, the more people try it and the greater its perceived business value. That’s why heavily cited architectures (e.g., the Transformer (Vaswani et al., 2017)) have become the backbone of many commercial systems. Beyond citation counts, the robustness of our findings is supported by analyses across multiple impact metrics. The predictive power of self-rankings held true not only for citations tracked by Semantic Scholar but also for Google Scholar citations and for a community engagement metric, GitHub stars. While these results are compelling, we acknowledge the need for validation over a longer time horizon and through experiments at more conferences to better assess the long-term impact of research work. To this end, we have already conducted similar experiments at ICML 2024 and 2025, as well as NeurIPS 2025, which we expect will validate and extend these findings. Further challenges include ensuring truthful reporting of self-rankings and determining how selfrankings can be formally incorporated into the review process. Notably, the comparative design of self-rankings renders it impossible for authors to simply inflate the quality of all their submissions. Indeed, an emerging body of literature is developing mechanisms in the game-theoretic framework that leverage rankings for more accurate review scores (Su, 2021; Yan et al., 2025). These results carry significant implications for the future of scholarly evaluation, particularly in fast-moving fields like AI. Such fields demand review systems capable of separating methodological correctness from likely scientific consequences. Current peer review in AI research tends to reward correctness and quantifiable improvements over broader impact. The increasing use of language models in the review process may further bias evaluations toward surface-level characteristics rather than profound scientific contributions.

Limitations. 6 Control for the confounding factors Several potential confounders have already been addressed in our analyses. Figures 2 and 3 stratified submissions by final decisions, showing that within both accepted and rejected papers, higherranked submissions accrued substantially more citations. Figure 4 further stratified by review scores, indicating that higher-ranked papers consistently received more citations even when controlling for similar scores. Our design also implicitly accounted for paper topics: for each author, the highest- and lowest-ranked papers were compared, thereby matching on author background and research area and (approximately) controlling for topic-related confounding. To further validate previous findings, we examine additional potential confounders in this section. Publication time is one potential factor, as earlier arXiv postings tend to accumulate more citations. The upper panel of Figure S.4 shows the posting-time distributions for high- and lowranked papers, with three key dates highlighted by red dashed lines. The distributions are similar, with mean posting dates of March 4, 2023 for high-ranked and March 8, 2023 for low-ranked papers, a difference of only four days. Another factor is the authors’ preferences. First, authors may self-cite high-ranked papers more frequently.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can AI systems perform peer review as effectively as humans? What human oversight must AI research systems have? How can we detect and account for LLM involvement in academic writing? Do restrictions on reviewer LLM use actually shape peer review behavior? Does AI-assisted research sacrifice exploration breadth for productivity gains?