Can institutional publication records train better scientific evaluators?
Can AI models learn to make reliable low-verifiability judgments by training on where and what fields published, rather than explicit quality rubrics? This matters because science depends on gatekeeping decisions that individual reviewers struggle to make consistently.
The authors argue that institutional traces, meaning the record of "what fields published, where, and at which tier," can train AI evaluators to make the low-verifiability judgments that science depends on. Across eight social science fields they built held-out four-tier research-pitch benchmarks and fine-tuned LLMs on field-specific publication outcomes. In management, the best single fine-tuned model (Qwen3-4B) reached 59.2%, 17.6 points above an expert majority vote of 41.6% from 48 gatekeepers and 28.1 points above the mean of 11 frontier reasoning models (31.1%). Best single-model accuracy ran from 55.0% in public administration to 85.5% in psychology. The fine-tuned models' confidence also rose on correct predictions and fell on wrong ones.
The excerpt says "individual reviewers are noisy, but fields repeatedly select, reject, and rank work through editorial decisions and journal hierarchies." Individual reviewers agree "at barely above chance" (kappa of 0.047 in the authors' own 48-expert panel, 0.17 in a meta-analysis of 48 studies), yet the aggregate system "deposits consistent quality stratification into the publication record." Fine-tuning reads that stratification back out. The evaluated models never see journal identity; each article was reduced to a research-pitch text, and training used a single tier-label token. The authors treat tier as "an institutional proxy for field-level evaluation, not as ground-truth quality." Frontier and base models stayed near chance in every field, including psychology, which the authors read as a transmission problem rather than a prompting problem.
This sharpens two neighboring notes. Can models learn what makes research worth doing? also aims at scientific taste, but supervises with citation counts. The authors call citations "a problematic proxy for evaluative judgment at the moment of decision," because they accumulate over years and are confounded by venue prestige and author networks; publication tier records the gatekeeping decision itself. The paper also gives a concrete route to the bottleneck named in What makes accountable judgment scarce when AI cognition is cheap?: when generation is cheap, the scarce step is deciding which candidate deserves work, and here that step is trained from institutional records rather than written as a rubric. Like Can readers tell truth from fabrication without evidence signals?, it places discernment in an explicit signal, saying the judgment "was never absent; it was simply never the explicit training target."
The excerpt does not establish as much as the headline suggests. Every benchmark is in the social sciences, and the authors call whether traces work in STEM "an empirical question." The human comparison covers management only; the authors say it "calibrates the magnitude rather than founding the mechanism." Tier boundaries were set by the authors' own domain experts, and the date screen is "a temporal-control measure rather than a guarantee of exclusion" from proprietary training data. The implication is that institutional records are a credible, low-cost training signal for evaluators in the social sciences, and that the claim about science generally awaits the STEM and human-benchmark work the authors name.
Inquiring lines that read this note 21
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems perform peer review as effectively as humans?- Can multi-stage AI review pipelines catch scientific flaws better than simple language models?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Could AI improve peer review rigor and catch human-missed errors?
- Why does publish-or-perish incentivize quantity over quality in research?
- What makes disruptive scientific work harder to publish and recognize?
- How can arXiv and journals scale quality control for AI-generated research?
- Can institutional statements alone correct misconceptions from unreviewed papers?
- What makes rhetorical polish misleading in evaluating research quality?
- Do citation counts better capture scientific quality than publication venue tiers?
- Why do individual peer reviewers show such low agreement on research merit?
- How does opaque AI methodology undermine peer review and reproducibility?
- Can technical accuracy in AI training data replace human review before publication?
- Where does AI assistance become reliable versus prone to failure in science?
- When should domain experts verify AI research claims before publication?
- Does institutional trace learning work equally well in STEM fields and social sciences?
- What distinguishes verifiable AI research domains from open-ended scientific questions?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can models learn what makes research worth doing?
Can large language models be trained to recognize high-impact research directions by learning from citation patterns? This explores whether 'scientific taste'—the judgment of what work matters—is a learnable skill separate from execution.
same goal of learning scientific taste; this paper supervises with publication tiers, not citation counts.
-
What makes accountable judgment scarce when AI cognition is cheap?
When AI systems can perform cognitive tasks cheaply and at scale, what human capabilities become most valuable? This explores whether judgment, verification, and accountability are the true bottlenecks in labor markets shaped by generative AI.
both treat evaluation rather than generation as the bottleneck; this paper supplies a training route to it.
-
Can readers tell truth from fabrication without evidence signals?
When readers see fluent text with no provenance information, do they distinguish accurate claims from AI-generated hallucinations? This tests whether presentation authority alone misleads judgment.
parallel claim that discernment appears only when an explicit signal is supplied.
-
Can machines learn to predict which research ideas will work?
Can a fine-tuned language model with access to published papers predict which unimplemented AI ideas will succeed empirically, and would it outperform human researchers making the same judgment?
evidence for: a fine-tuned GPT-4.1 with paper retrieval beats expert NLP researchers 64.4 to 48.9 percent, while frontier models sit at chance
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLMs learn scientific taste from institutional traces across the social sciences
- Stop Automating Peer Review Without Rigorous Evaluation
- LLM-REVal: Can We Trust LLM Reviewers Yet?
- How to Find Fantastic AI Papers: Self-Rankings as a Powerful Predictor of Scientific Impact Beyond Peer Review
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
- Predicting Empirical AI Research Outcomes with Language Models
- The Emerging AI Paper-Review Arms Race: Adversarial Co-Evolution in Scholarly Publishing
Original note title
fine-tuning on publication-tier outcomes yields evaluators that beat frontier models and experts — management 59.2% against 41.6% expert majority