Does a paper's citation count say more about whether its science is good than the prestige of the journal that printed it?
Do citation counts better capture scientific quality than publication venue tiers?
This explores whether counting how often a paper gets cited is a better signal of good science than where it was published (top journal vs. lesser venue). The collection has no head-to-head test, but it does treat both signals as training data, and it shows how each one can mislead.
This explores whether citation counts or venue prestige better tell you a paper is good. The short answer: the collection doesn't settle this directly. What it shows is more interesting. Each signal holds real information that models can learn from, and each can be distorted in its own way.
Start with the case for each. One group fine-tuned models on where social science research pitches ended up being published. Those models beat both frontier reasoning models and majorities of expert reviewers at predicting outcomes, reaching 59.2% accuracy in management, where experts agreed only 41.6% of the time Can institutional publication records train better scientific evaluators?. The models learned a field's unwritten standards just from which journals accepted what. A separate project took the citation route. It trained models on 700,000 pairs of papers matched on citations, and they learned a kind of 'scientific taste': they predicted impact better than GPT-5.2 and came up with higher-impact ideas Can models learn what makes research worth doing?. So venue tiers capture what gatekeepers value, and citations capture what the community later uses. Those are two different things, and both can be learned.
Now the cracks. Citations can grow for reasons that have little to do with quality. Researchers who use AI publish about 3× more papers and get 4.8× more citations. Over the same period, science as a whole covers fewer topics and has less collaboration, because work piles up on data-rich problems Does AI help individual scientists while narrowing scientific focus?. A high citation count can reward following the crowd. A finding from a different setting suggests why counts carry weight anyway: users prefer AI answers with more citations almost as much when the citations are irrelevant Do users trust citations more when there are simply more of them?. People treat volume as credibility.
Venue tiers have their own weak spot: the peer review that produces them. One model shows how a flood of submissions overloads unpaid reviewers. Review accuracy drops, authors respond by submitting more speculative work, and submissions climb further Does peer review quality collapse under submission overload?. A survey of 230 publications describes an arms race. AI scales up paper production, AI-assisted review tries to keep pace, and manipulation and countermeasures escalate together Does AI create a coupled arms race in research production and review?. Conferences are now desk-rejecting papers with made-up references How can conferences detect and handle LLM misuse in peer review?. Meanwhile, some work goes around venues entirely. One arXiv preprint shaped public debate long before MIT said it had no confidence in it Can unreviewed preprints shape scientific debate before peer review?.
The surprising takeaway is that both signals are now being turned into AI evaluators. Whatever bias a signal has, the evaluator trained on it will inherit. A model trained on citations may learn to favor crowded, data-rich topics, and a model trained on venues may learn to favor whatever an overloaded review system lets through. The better question may be which kind of distortion you're willing to automate.
Sources 8 notes
LLMs fine-tuned on eight social science publication records beat both expert majority votes and frontier reasoning models at evaluating research pitches, reaching 59.2% accuracy in management versus 41.6% expert agreement. The models learned field-level evaluation logic from institutional stratification rather than written criteria.
Reinforcement learning trained on 700K citation-matched paper pairs successfully teaches models to predict research impact better than GPT-5.2 and generate higher-impact research ideas. Scientific taste emerges as a community-aligned capability distinct from execution skills.
AI-augmented researchers publish 3× more papers and receive 4.8× more citations, but collective science shrinks topic coverage by 4.63% and researcher collaboration by 22%. AI concentrates work on data-rich problems rather than exploring new questions.
Analysis of 24,000 Search Arena interactions shows irrelevant citations boost user preference (β=0.273) nearly as much as relevant citations (β=0.285), indicating citation count functions as a decoupled trust heuristic.
A two-journal model shows that rising submissions overtax unpaid reviewers, forcing journals to recruit less qualified reviewers or overload existing ones, which drops review accuracy and incentivizes authors to submit more speculatively, driving submissions higher. The mechanism is structural but its empirical strength remains to be measured.
Show all 8 sources
A survey of 230 publications reveals production scaling, evaluation automation, manipulation, defenses, evasion, and ecosystem feedback as linked response relations among actors. Evidence is strongest for early stages and weakens toward long-horizon adaptation and feedback.
Program chairs used imperfect detectors as one input for area chairs rather than automated filters, but desk-rejected papers with confirmed fabricated references as a tractable enforcement point. Multiple human review steps mitigated false positives.
MIT's case demonstrates that an arXiv preprint shaped AI and science discussions extensively despite never undergoing peer review. When the institution later raised reliability concerns, the damage to discourse had already occurred.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Stop Automating Peer Review Without Rigorous Evaluation
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
- LLMs learn scientific taste from institutional traces across the social sciences
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- Artificial Intelligence Tools Expand Scientists' Impact but Contract Science's Focus (Just accepted by Nature, to be online soon)
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- Evaluating Sakana's AI Scientist: Bold Claims, Mixed Results, and a Promising Future?