Can authors rank their own papers better than peer reviewers?
Do researchers have better insight into which of their own submissions will prove influential than official peer reviewers do? This matters because peer review is expensive and may miss work with long-term scientific value.
At ICML 2023, the paper's team asked researchers with several submissions to rank them by perceived quality, through a platform they built (OpenRank.cc) with the conference organizers' approval. The paper reports that 1,342 researchers ranked 2,592 submissions, and that the papers each author ranked highest drew, on average, twice the citations of the ones ranked lowest, for accepted and rejected papers alike. Of the 22 papers with over 150 citations, 17 had been ranked highest by their authors. The sharpest comparison is that self-rankings "outperformed peer review scores in predicting future citation counts."
The paper grounds this in who knows the work best. Authors "possess unique understanding of their work's conceptual depth and long-term promise," while overloaded reviewers "may prioritize quantifiable gains over a submission's broader scientific implications." The comparative format matters because authors "cannot simply claim all their submissions are of the highest quality." Because each author's highest and lowest papers are compared, the design also roughly matches on author background and research area. The paper cites game-theoretic work (Su, 2021, 2025; Yan et al., 2025) for the claim that this format can incentivize truthful reporting under certain conditions.
Set against the nearest notes, this asks a different question from Does banning LLM use in peer review change review outcomes?, which measures whether review policies move scores. This paper asks which signal review should draw on at all. It also gives empirical footing to the warning in How much does rhetorical style shift AI review scores?. The discussion says LLM use "may further bias evaluations toward surface-level characteristics rather than profound scientific contributions," and the rhetoric study shows LLM scores moving with presentation. The contrast with Can inference scaling help reviewers catch errors humans miss? is sharper. That paper searches manuscripts for correctness failures, while this one separates "methodological correctness" from "likely scientific consequences." A reviewer tuned to the first may not capture the second.
The excerpt leaves much unestablished. Citations are the paper's own proxy for impact, which it calls "widely used, albeit imperfect." The evidence comes from one conference and one survey year. The paper says similar experiments at ICML 2024 and 2025 and NeurIPS 2025 have already been conducted, but it reports none of their results. The citation analysis covers 1,527 submissions, not 2,592, after keeping only exact title-and-author matches on Semantic Scholar, and the survey response rate was 30.4%. The excerpt gives the preprint check (mean posting dates of March 4 and March 8, 2023) but cuts off before the self-citation result, and it reports the Google Scholar and GitHub-star results without figures. The defensible implication is that a self-ranking step is worth testing as a complement to review scores for finding influential AI work. The excerpt does not show that self-rankings should replace review, or that they track quality beyond citations.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does banning LLM use in peer review change review outcomes?
Can policies restricting or allowing AI tools shift how reviewers score papers and make decisions? This matters because review quality and fairness depend on consistent standards.
another peer-review experiment; asks whether review rules move outcomes, not which signal to use
-
How much does rhetorical style shift AI review scores?
When manuscripts are rewritten to improve rhetoric while keeping scientific content identical, do LLM reviewers change their scores? Understanding this matters for ensuring AI-assisted peer review evaluates substance, not polish.
supplies evidence for the source's worry that LLM review drifts toward surface characteristics
-
Can inference scaling help reviewers catch errors humans miss?
Explores whether spending extra compute at review time—checking proofs and experiments line by line—can surface deep flaws that evade human expert reviewers, and how this scales with AI-assisted submissions.
contrast: correctness checking versus the impact signal self-rankings target
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- Screening, sorting, and the feedback cycles that imperil peer review
- Stop Automating Peer Review Without Rigorous Evaluation
- AI Can Learn Scientific Taste
- Do LLMs Favor LLMs? Quantifying Interaction Effects in Peer Review
- Can large language models provide useful feedback on research papers? A large-scale empirical analysis
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- More Versus Better, Part I
Original note title
authors' self-rankings predicted citations better than ICML review scores did — top-ranked papers drew twice the citations of bottom-ranked ones