LLMs learn scientific taste from institutional traces across the social sciences
Reinforcement-learned reasoning has powered recent AI leaps on verifiable tasks — mathematics, code, structure prediction. The harder bottleneck is evaluative judgment in low-verifiability domains, where no oracle anchors reward and which untested ideas deserve attention is the question. We test whether institutional traces — the record of what fields published, where, and at which tier — can serve as a training signal for AI evaluators. Across eight social science disciplines (psychology, economics, communication, sociology, political science, management, business and finance, public administration), we built held-out four-tier research-pitch benchmarks and fine-tuned LLMs on field-specific publication outcomes. The fine-tuned models cleared the 25% chance baseline and exceeded frontier-model performance by wide margins, with best single-model accuracy ranging from 55.0% in public administration to 85.5% in psychology. In management — evaluated against 48 expert gatekeepers, 174 junior researchers, and 11 frontier reasoning models — the best single fine-tuned model (Qwen3-4B) reached 59.2%, 17.6 percentage points above expert majority vote (41.6%, non-tied) and 28.1 above the frontier mean (31.1%). The fine-tuned models also showed calibrated confidence: confidence rose when predictions were correct and fell when wrong, mirroring how a skilled reviewer can say “I’m sure” versus “I’m guessing”. Selective triage on this signal reached very high accuracy on the highest-confidence subsets in every field. Reinforcement learning with chain-of-thought plateaued below direct fine-tuning in the management mechanism probe, and a contemporaneous GPT-5.5 comparison found no accuracy gain from high-reasoning inference over chat/log-probability classification across all eight fields. The approach offers a cost-effective way to train evaluative models from institutional records. Institutional traces, we conclude, encode a scalable training signal for the low-verifiability judgment on which science depends.
Introduction. Artificial intelligence has advanced fastest in scientific domains where candidate outputs can be checked. Protein structures can be matched against experimental constraints[1]; mathematical proofs can be verified by formal systems[2]; computer code can be executed against tests[3]. The remarkable AI-for-science results of the last five years all share this property: an unambiguous oracle exists, and the model’s job is to find an output the oracle accepts.
Most of science, however, depends on a prior judgment with no such oracle. Before any experiment, proof, or benchmark exists, scientists must decide which untested ideas deserve scarce attention. This evaluative judgment, the capacity sometimes called taste, is what editors, reviewers, hiring committees, and grant panels enact whenever they choose what to pursue and what to set aside. As AI systems generate hypotheses, analyses, and manuscripts at unprecedented scale[4, 5], the bottleneck in science is shifting from production toward evaluation[6]: the rate-limiting step is no longer writing the next plausible idea, but deciding which plausible idea is worth working on. Scientific taste defines the next critical boundary in artificial intelligence, where capability drops sharply from domains with clear answers to those requiring discrimination and judgment under uncertainty[7].
Evaluative judgment is a paradigmatic case of collective tacit knowledge: understanding embedded in institutional practices that no individual participant fully articulates, yet that the system reliably enacts[8, 9]. Individual peer reviewers agree on quality assignments at barely above chance (meta-analytic Fleiss’ kappa of 0.17 across 48 studies[10]), and neither career stage nor editorial experience improves this reliability[11, 12]. Ethnographic evidence confirms that academic judgment operates through intuitive assessment and disciplinary sensibility, not rule application[13]. Yet the same institutional system, integrating thousands of such noisy judgments over decades, produces consistent quality stratification across publication tiers[14]. That apparent paradox, agreement-poor at the individual level and signal-rich at the institutional level, is the central observation that motivates this work.
Frontier language models confront the same boundary. They can summarize, reason, and write fluent reviews, but prompting them with evaluative criteria does not by itself produce selective judgment: prior AI-review studies document bias and blind spots, and the present evaluations show frontier models clustering near chance, over-predicting favorable categories, and compressing judgments into middle tiers[15, 16, 17, 18]. The results below argue against a prompt-engineering-only explanation and point instead to a transmission problem. Tacit standards do not survive the channel of explicit instruction, and reinforcement learning from human preferences pushes models toward agreeable, lenient assessments[19, 20].
We propose institutional traces (the historical record of what fields published, where, and in which prestige tier) as an alignment signal for AI evaluators when no verifiable reward is available. Individual reviewers are noisy, but fields repeatedly select, reject, and rank work through editorial decisions and journal hierarchies. These records are imperfect proxies for quality, but they preserve a learnable trace of field-level evaluative practice. If a model can be aligned to those traces, then low-verifiability scientific judgment may be made computationally accessible without first reducing it to a written rubric. Figure 1 summarizes this conceptual frame.
This paper is empirically novel against an active literature. Recent work has used reinforcement learning from verifiable reward to push frontier reasoning on math, code, and other check-able tasks[21]; specialized continued pretraining has produced LLMs that surpass human experts at predicting empirical neuroscience outcomes[22]; LLMs have been shown to generate research ideas judged more novel than expert-generated ones[4], to produce reviewer-quality qualitative feedback on Nature-family submissions[23], and to power end-to-end automated research pipelines that can complete AI-research workflows and pass workshop peer review[24, 25]. Concurrent work by Tong et al.[26] also pursues “scientific taste” through a related framing, using reinforcement learning from community feedback with citation counts on accepted papers as the supervision signal. Citations are a problematic proxy for evaluative judgment at the moment of decision: they accumulate over years, are confounded by venue prestige and author networks, and are not part of what reviewers see when deciding whether an idea is worth pursuing. Institutional selection traces (which venues accepted which work in the first place) capture the prior gatekeeping decision itself, which is the variable evaluators are actually computing.
Method. Benchmarks. We constructed balanced four-tier evaluation benchmarks in eight social science fields: management (organizational psychology and management), economics, business and finance, communication, political science, psychology, public administration, and sociology. Field scope follows Web of Science discipline categories. All source articles were published after 30 June 2025 and all benchmark items were excluded from SFT training; because frontier providers do not disclose complete training corpora, this date screen is a temporal-control measure rather than a guarantee of exclusion from every proprietary model. Management contained 120 pitches (30 per tier) drawn from a 19-journal source universe; the smaller size reflects the recruitment ceiling of the matched expert panel. The remaining seven fields each contained 200 pitches (50 per tier) drawn from field-specific journal universes. Benchmark tiers (exceptional, strong, fair, limited) are properties of the source journal, assigned by a human pre-construction procedure that is logically separate from the AI evaluation. In management, the management subject-matter experts on the author team directly assigned each of the 19 source journals to a tier — no external numerical proxy was used, since the management hierarchy is sufficiently codified in tenure and editorial-board norms — consulting additional field experts on boundary cases. In the other seven fields, field experts nominated candidate journals for each tier (drawing on editorial standing, tenure-credit norms, and submission/acceptance patterns), and the nominations were cross-checked against an external benchmark of citation impact metrics and recognized journal-quality lists used by tenure committees and field associations; inconsistencies between expert nomination and the external benchmark were resolved through additional expert consultation, iterated until convergence. The author team’s domain experts retained final decision authority on all tier boundaries. Crucially, AI models being evaluated never see journal identity — they predict tier from research-pitch text alone (full per-field mappings in SI Tables ST3–ST10). Each source article was transformed into a standardized research-pitch text using a fixed extraction workflow that exposes only the core research question and theoretical framing; methods, empirical results, journal identity, and author identity were removed, to isolate idealevel assessment across all evaluator classes. Throughout, “publication tier” is treated as an institutional proxy for field-level evaluation, not as ground-truth quality. Full benchmark construction, Web of Science scoping, and journal-mapping detail are in SI SM5.
Supervised fine-tuning. In management, where we conducted the full human, frontier-model, mechanism, and transfer analyses, we fine-tuned four architectures: GPT-4.1, GPT-4.1-nano, Qwen3-30B- A3B-Instruct (30B-parameter mixture-of-experts, 3B active at inference), and Qwen3-4B-Instruct (4Bparameter dense). The four management SFT models occupied a narrow in-domain accuracy band (55.0– 59.2%), while GPT-4.1 accounted for most external fine-tuning spend; we therefore used the three costeffective architectures (Qwen3-30B-A3B, Qwen3-4B, and GPT-4.1-nano) for the full eight-field core. GPT-4.1 was retained only for management mechanism probes (pairwise discrimination and inputcompression transfer) and the cross-field-transfer test, where larger representational capacity was expected to matter. Architecture-matched un-fine-tuned base controls were evaluated on every field’s benchmark to isolate the SFT effect. Each training example paired a research-pitch text, presented through a prespecified field-agnostic evaluation prompt, with a single tier-label completion token. Training minimized label-token negative log-likelihood with input tokens masked from gradient updates, forcing the model to learn the mapping from research content to quality tier. Training corpora were field-specific and disjoint from evaluation benchmarks: management used 4,479 pitch–outcome pairs; the seven other fields used 2,094–5,593 pairs each (per-field sizes in SI Table ST11).
Discussion. Scientific taste, the capacity to judge which untested ideas are worth pursuing, has been treated as an irreducibly individual property[8, 13]. The data here reframe it as an institutional property. Across eight social science disciplines, fine-tuning on publication records produced evaluative discrimination that neither frontier AI systems nor experienced human gatekeepers achieved through reasoning or expertise alone. The signal was never absent; it was simply never the explicit training target. Individual reviewers agree at barely above chance (kappa = 0.047 in our 48-expert management panel; meta-analytic kappa = 0.17 across 48 prior studies[10]), yet the aggregate institutional system deposits consistent quality stratification into the publication record[14]. Fine-tuning extracts what peer-review committees collectively encode but no individual reviewer can articulate: Collins’s collective tacit knowledge[9] made computationally accessible.
Cross-field variation is graded but universal in direction. Psychology yields the highest fine-tuned accuracy (85.5%) and public administration the lowest (55.0%), and the broad high-to-low gradient is stable across architectures even though exact rankings vary; exploratory cross-field transfer of management-trained checkpoints is bounded by model scale rather than by surface field proximity. Crucially, frontier and base models cluster near chance in every field, including psychology — Gemini 3.1 Pro reaches 33.5%, GPT-5.5 High reaches 29.5%, and GPT-5.5 chat/log-probability reaches 31.5% in the field where fine-tuning reaches 85.5%. Pretraining alone, even on the journal-prestige hierarchies themselves, does not differentiate easy from hard fields. The easy/hard gradient is only legible to a model that has been given explicit access to institutional outcomes; surface text does not reveal it.
The GPT-5.5 evaluation also tests a stronger version of the “frontier models will catch up” objection. GPT-5.5 chat/log-probability was near chance and poorly calibrated across the eight benchmarks, and in that evaluation GPT-5.5 High reasoning was no better than GPT-5.5 chat. This pattern matters because it separates general model progress from the missing training signal. More capable models can improve fluency, breadth, and tool use, but evaluative judgment over low-verifiability ideas still requires exposure to the institutional outcomes that define the field’s tacit standard.
If institutional taste can be recovered cost-effectively from publication records, the practical implication is upstream: AI assistance shifts from generation to evaluation. Frontier systems and automated-research agents can now generate hypotheses, analyses, and complete manuscripts at speed[24, 25, 34]; the harder bottleneck is the reward signal needed to decide which candidate deserves further work. In verifiable domains, that signal can come from tests, proofs, or experiments. In low-verifiability domains, institutional traces offer a way to train the evaluator itself. Several properties of fine-tuned models make this triage role realistic in the present setting: calibrated confidence in every field tested (mean top-10% accuracy 94.6%), pairwise transfer to formats absent from training (84.3% in management), cross-architecture consensus that concentrates reliable predictions into a high-precision subset, adjacent-tier rather than arbitrary residual errors, and management temporal stability sufficient that a model trained on decade-old decisions still beats frontier systems and expert panels.
Limitations. The study deliberately combines eight-field breadth with management-only depth: senior gatekeeper recruitment was tractable for the authors only in management, so we used that field for the deepest probes while validating the institutional-trace phenomenon broadly with cost-effective models elsewhere. The within-architecture fine-tuned-versus-base comparison, universally positive across all 24 model–field combinations, identifies institutional traces as the operative training signal without requiring a human comparator in every field. The management human benchmark calibrates the magnitude rather than founding the mechanism. Extending human benchmarks to psychology, where fine-tuned accuracy is highest, and public administration, where it is lowest, would further test the boundaries of institutionaltaste learning, and is the most natural next experiment.
Several substantive caveats remain. The present benchmarks are confined to the social sciences; whether institutional traces work similarly in STEM is an empirical question, especially because reproducibility, citation dynamics, and venue hierarchies play different roles across STEM fields.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can AI systems perform peer review as effectively as humans?- Can multi-stage AI review pipelines catch scientific flaws better than simple language models?
- Can workshop acceptance rates reliably measure AI research quality compared to main conferences?
- Could AI improve peer review rigor and catch human-missed errors?
- Why does publish-or-perish incentivize quantity over quality in research?
- What makes disruptive scientific work harder to publish and recognize?
- How can arXiv and journals scale quality control for AI-generated research?
- Can institutional statements alone correct misconceptions from unreviewed papers?
- What makes rhetorical polish misleading in evaluating research quality?
- Do citation counts better capture scientific quality than publication venue tiers?
- Why do individual peer reviewers show such low agreement on research merit?
- How does opaque AI methodology undermine peer review and reproducibility?
- Can technical accuracy in AI training data replace human review before publication?
- Why does automated evaluation consistently overestimate research quality?
- Where does AI assistance become reliable versus prone to failure in science?
- When should domain experts verify AI research claims before publication?
- Does institutional trace learning work equally well in STEM fields and social sciences?
- What distinguishes verifiable AI research domains from open-ended scientific questions?
- Where do human researchers retain competitive advantage over autoresearch systems?