AI learns procedures well, but only from what someone wrote down, so can it ever pick up the know-how labs pass on by doing?
Can AI models learn tacit procedural knowledge that exists only in laboratory practice?
This explores whether AI can pick up the unwritten, hands-on know-how of lab work (the feel for a protocol that researchers pass along by doing, not by writing), and what the corpus says about how models acquire any kind of 'how-to' knowledge at all.
This explores whether AI can learn the unwritten, hands-on know-how of lab work: the things researchers learn at the bench and rarely put on paper. The short answer is that the corpus says nothing directly about wet-lab practice. What it does have is a clear picture of where models get their 'how-to' knowledge, and that picture explains why lab know-how is so hard to capture. The surprising part is that models are already good at procedural knowledge. They just get it from text. An analysis of 5 million pretraining documents found that model reasoning draws on broad, reusable procedures spread across many sources, while factual recall depends on memorizing specific documents Does procedural knowledge drive reasoning more than factual retrieval?. So models can learn procedures well, but only from procedures someone wrote down.
That sets a hard limit. If a technique lives only in a senior postdoc's hands, it isn't in the training data. Clever prompting can't help: prompt optimization can only bring out knowledge a model already has and can't add knowledge that's missing Can prompt optimization teach models knowledge they lack?. Several papers show the same pattern from another direction: post-training mostly brings out abilities that were already in the base model and creates few new ones Do base models already contain hidden reasoning ability?. For lab know-how, that means more training on the existing literature mostly brings out what the literature already contains.
The word 'tacit' can mislead here. Philosophers have argued that LLMs may have tacit knowledge in Martin Davies' sense: knowledge built into their internal structure that they can't state outright. Some early evidence comes from experiments that edit a single stored fact inside a model Do language models possess tacit knowledge in Davies' sense?. But that is tacit in the sense of unspoken inside the model. It isn't the same as the embodied, tacit-in-the-world skill of knowing when a cell culture looks wrong. A model can have the first kind without ever touching the second.
The more promising doorways are the ones that change the learning signal itself. Training on expert demonstrations alone caps an agent at what the curators thought to record, and the agent never learns from its own failures Can agents learn beyond what their training data shows?. But demonstrations can also be used to infer the hidden standards experts work by. One method learns what 'good' looks like from expert examples in fields that have no automatic answer-checker Can reasoning emerge from expert demonstrations alone?. That is close to how apprentices learn. Other work points toward learning by doing. Systems improve through trial, error and testing against real results rather than by reading Can AI systems improve themselves through trial and error?. Models trained to predict what happens next in an environment can also stand in for real practice Can language models learn to simulate agent environments?. Put together, the corpus suggests lab know-how won't come from reading more papers. It might come from recording practice as it happens (demonstrations, logs of what was tried, observed outcomes) and letting models learn from their own mistakes against that record. Whether this works for physical lab skills is still an open question the collection doesn't answer.
Sources 8 notes
Analysis of 5 million pretraining documents shows reasoning relies on broad, transferable procedural knowledge from diverse sources, unlike factual recall which depends on narrow, document-specific memorization of target facts.
Prompting works entirely within a model's pre-existing training distribution and cannot supply domain knowledge absent from training data. This creates a hard ceiling: no prompt strategy can compensate for missing foundational knowledge, only reorganize what already exists.
Five independent mechanisms—RL steering, critique fine-tuning, decoding changes, SAE feature steering, and RLVR—all elicit reasoning already present in base model activations. Post-training selects rather than creates reasoning; the bottleneck is elicitation, not capability acquisition.
Transformer LLMs can meet Davies' criteria for tacit knowledge based on architectural features and causal tracing with ROME edits. However, evidence rests on a single fact-editing case, and replication challenges suggest the causal localization may not be as precise as initially claimed.
Agents trained on static expert datasets cannot learn from their own failures or generalize beyond demonstrated scenarios because they never interact with environments during training. Competence is capped by what curators imagined, not by agent capacity.
Show all 8 sources
RARO recovers implicit reward functions from expert demonstrations through adversarial co-training between a reasoning policy and relativistic critic. This approach matches verifier-based RL performance on reasoning tasks while extending to domains lacking automated verification.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
Qwen-AgentWorld demonstrates that native language world models trained via next-state prediction on 10M+ trajectories outperform real-environment training on three benchmarks and transfer across seven domains, positioning next-state prediction as a foundation objective for agents.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Eliciting Reasoning in Language Models with Cognitive Tools
- Agent Learning via Early Experience
- On the Interplay of Pre-Training, Mid-Training, and RL on Reasoning Language Models
- The Darwin Gödel Machine: AI that improves itself by rewriting its own code
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Qwen-AgentWorld: Language World Models for General Agents
- What Do Large Language Models Know? Tacit Knowledge as a Potential Causal-Explanatory Structure
- Hyperagents