Scientific consensus isn't just a pile of correct claims; it's a social achievement, built on unwritten judgment that a community learns to trust.
What role does tacit knowledge play in expert consensus on frontier science?
This explores how unwritten, hard-to-articulate know-how (the judgment experts carry but can't fully write down) shapes what scientific communities agree on at the cutting edge, and whether AI can share in it. The corpus has no note on tacit knowledge in consensus-building itself, so this answer pieces it together from nearby material.
This explores how unwritten, hard-to-articulate know-how shapes what scientific communities agree on at the frontier, and whether AI can take part. Be warned up front: the collection has no note that studies tacit knowledge in consensus formation directly. What it does have are several notes that circle the same territory from different sides, and together they give a fairly sharp picture.
The strongest thread is that expert consensus isn't just a sum of correct claims. It's a social achievement. Can AI ever gain expert community trust through participation? argues that experts earn authority through participation, track record and standing inside a community, and that this is exactly where tacit judgment gets tested and trusted. Other experts learn whose instincts hold up. Can language models distinguish expert arguments from common assumptions? makes the matching point from the reader's side: an argument's weight depends partly on who makes it, and language models, working only from text, lose that signal. They can't easily tell a seasoned expert's hunch from a common assumption phrased the same way. So on this view, tacit knowledge matters for consensus less as private know-how and more as a reputation for judgment that the community has learned to read.
The surprise is that "tacit knowledge" means something quite different in the AI literature. Do language models possess tacit knowledge in Davies' sense? uses a philosopher's sense: knowledge that is encoded in a system's internal workings without being stated as explicit rules. In that sense LLMs might have it, though the evidence is thin and rests on one fact-editing study. Read next to Can LLMs predict novel scientific results better than experts?, where fine-tuned models beat neuroscientists at predicting which experimental results actually happened, it starts to look like models may carry something like an expert's unspoken sense of what is plausible. That sense is what makes them good forecasters, and it's also what makes them hallucinate. Still, a model that predicts well hasn't joined the community, and the first two notes say that membership is what turns judgment into consensus.
The remaining notes show what happens when that social filter comes under strain. Does AI-generated knowledge have the same structure as hearsay? and Is AI returning knowledge to flow-based economies? both argue that AI output lacks an accountable carrier, a person whose standing backs the claim, so the traditional checks on knowledge have nothing to hold on to. Can AI generate knowledge faster than humans can evaluate it? adds that the experts whose judgment could sort this material don't scale. Can unreviewed preprints shape scientific debate before peer review? shows a concrete case: an unreviewed preprint shaped the debate before anyone with standing had vetted it. On the research side, Do frontier AI agents actually conduct novel research or just optimize? finds that frontier agents mostly recombine known techniques rather than make the kind of new judgment calls experts are trusted to make.
The takeaway you may not have expected: the debate over whether AI "has" tacit knowledge may be the wrong question for science. Even if models carry implicit know-how inside them, frontier consensus depends on tacit judgment that other experts can see, test and vouch for over time. That accountability layer is what AI currently lacks, and it's the layer the flood of AI-generated papers is wearing down.
Sources 9 notes
Expertise is validated through social participation and track record within expert communities, not individual accuracy alone. AI cannot enter this validation circle because it lacks social embeddedness, testable judgment history, and ability to participate in the consensus-building processes that define expert paradigms.
LLMs lose the social context that gives expert claims their force—reputation, track record, and standing—because they process only text, not the social world where expertise is built and evaluated.
Transformer LLMs can meet Davies' criteria for tacit knowledge based on architectural features and causal tracing with ROME edits. However, evidence rests on a single fact-editing case, and replication challenges suggest the causal localization may not be as precise as initially claimed.
BrainBench benchmarks show fine-tuned LLMs outperform neuroscience experts at predicting which experimental results actually occurred. The same pattern-integration tendency that causes hallucination in retrieval tasks enables genuine prediction in forward-looking scenarios.
AI output shares all defining features of hearsay: testimony at remove, modification in retelling, unattributable origin, and unverifiability against stable sources. This means Enlightenment verification tools—citation, archiving, peer review, evidentiary chains—cannot process AI output by design.
Show all 9 sources
Print culture fixed knowledge as accumulated stock; AI returns knowledge to generative flow. However, unlike oral and gift economies, AI flows lack the embodied transmission—the speaker, the giver—that historically anchored knowledge circulation.
AI produces knowledge faster than human judgment can verify it, collapsing epistemic confidence just as monetary hyperinflation collapses purchasing power. The gap self-reinforces because evaluation tools are themselves AI-generated, trapping the system in acceleration.
MIT's case demonstrates that an arXiv preprint shaped AI and science discussions extensively despite never undergoing peer review. When the institution later raised reliability concerns, the damage to discourse had already occurred.
Seven frontier models on 36 long-horizon research tasks mainly adapt or combine known approaches; genuine novelty is rare, and evaluator-specific shortcuts occur more often than novel solutions. Performance varies substantially across runs.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Epistemic Deference to AI
- "That's AI Slop, You Bot!" Studying Accusations, Evidence, and Credibility in Online Discourse Towards LLM-Generated Comments
- A Rational Analysis of the Effects of Sycophantic AI
- Six misconceptions about large language models: A minimal model and diagnostic taxonomy
- Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
- Stop Automating Peer Review Without Rigorous Evaluation
- What Do Large Language Models Know? Tacit Knowledge as a Potential Causal-Explanatory Structure
- Large language models surpass human experts in predicting neuroscience results