INQUIRING LINE

Why does teaching an AI the hidden thinking behind expert writing help it reason well in totally new fields?

Why does reconstructing hidden thought processes during training improve cross-domain transfer?

This explores why training a model on expert writing plus a reconstruction of the thinking behind it (self-talk, recalling facts, checking work) produces reasoning that carries over to new subject areas, when training on the finished text alone does not.


This explores why adding back the thinking behind expert writing helps a model reason in fields it wasn't trained on. The key idea is that a finished expert text is a residue. A proof, a legal brief or a clinical note shows the conclusions and hides the self-talk, the recalled background knowledge and the checks the expert ran along the way. When researchers rebuild those hidden steps and train on the text and its reconstructed thinking together, the resulting reasoning skills transfer across domains. The model also learns to think longer on hard problems and shorter on easy ones, and it beats standard continued pretraining by up to 8 points on the hardest problems Can reconstructing expert thinking improve reasoning transfer?. The likely reason for the transfer is that the subject matter changes between fields while the thinking moves stay the same. Recalling what you know, checking a step and noticing you're stuck look much the same in chemistry and in law. Training on finished text teaches what experts conclude. Training on the reconstruction teaches how they get there, and that part travels.

A useful parallel comes from an unrelated corner of the collection, recommendation systems. There, turning an item's text description into discrete codes, and only then into embeddings, transfers across product domains better than encoding the raw text directly. The intermediate step strips away surface-level text bias Can discrete codes transfer better than text embeddings?. Reconstructed thinking may do the same job for reasoning. It gives the model a layer of process that sits between a domain's vocabulary and its conclusions, so the model learns less of the field's surface style and more of its structure. A related result points the same way. Rewarding a model for coherent explanations, and not only for matching the right tokens, builds domain knowledge more deeply than standard supervised fine-tuning Can reinforcement learning embed domain knowledge more effectively than supervised fine-tuning?. In both cases, making the reasoning itself the training target is what helps.

Another part of the collection reframes what this training is actually doing. Several independent lines of evidence suggest base models already hold reasoning ability that minimal training simply brings out. On that view, post-training selects reasoning more than it creates it Do base models already contain hidden reasoning ability?. If that's right, reconstructed thinking may work less by teaching new skills and more by showing the model, early and across many domains, when to use skills it already has. That would also explain why the gains appear during pretraining. A separate approach treats chain-of-thought as an exploratory move during pretraining itself. It rewards any reasoning that makes the following text easier to predict, which needs no external answer-checker, and it lifts math and science scores substantially Can chain-of-thought reasoning be learned during pretraining itself?. Both results suggest reasoning can be planted earlier in training than the usual fine-tuning-afterwards approach assumes.

The thinking doesn't always have to be written out in words. Some models reason through repeated computation in their hidden internal states, without producing any intermediate text. A 150M-parameter model built this way broke the cost-versus-accuracy frontier on the ARC-AGI-1 puzzle benchmark Can latent reasoning match chain-of-thought cost efficiency without verbalizing?. Researchers can also detect what a model is poised to say internally before it says it, which suggests that internal thinking has real structure Can we read a language model's unspoken thoughts?. Written-out reconstructed thoughts may be one way of training that internal workspace, and not the only one. There's also a caution about how far any single result generalizes. Gains from reasoning data depend heavily on the base model, the answer-checker, the optimizer and the training budget, so the same data can have different effects in a different setup What is the actual reusable unit of reasoning data?. The collection has one direct study of reconstructed expert thinking. The explanation for why it transfers is pieced together from the neighboring work above, and no single paper tests it directly.


Sources 8 notes

Can reconstructing expert thinking improve reasoning transfer?

Training on expert texts augmented with reconstructed thought processes (self-talk, knowledge recall, verification) produces reasoning skills that transfer across domains and adapt depth to problem difficulty, outperforming standard continual pretraining by up to 8 points on hard problems.

Can discrete codes transfer better than text embeddings?

VQ-Rec demonstrates that mapping item text to discrete codes via product quantization, then to embeddings, improves cross-domain transfer compared to direct text encoding. The discrete intermediate reduces text bias and enables efficient per-domain fine-tuning.

Can reinforcement learning embed domain knowledge more effectively than supervised fine-tuning?

RLAG rewards both answer accuracy and explanation rationality by cycling between augmented and unaugmented generation, progressively internalizing coherent knowledge structures. This outperforms SFT because it prioritizes reasoning quality over token-level correctness.

Do base models already contain hidden reasoning ability?

Five independent mechanisms—RL steering, critique fine-tuning, decoding changes, SAE feature steering, and RLVR—all elicit reasoning already present in base model activations. Post-training selects rather than creates reasoning; the bottleneck is elicitation, not capability acquisition.

Can chain-of-thought reasoning be learned during pretraining itself?

RLP treats CoT as exploratory action during pretraining, using log-likelihood improvement as verifier-free reward. Applied to Qwen3-1.7B and Nemotron-Nano-12B, the method improves math and science benchmarks substantially, suggesting reasoning can be planted earlier in training.

Show all 8 sources
Can latent reasoning match chain-of-thought cost efficiency without verbalizing?

A 150M-parameter model combining in-context demonstrations with iterative latent computation reached 29.5% pass@2 on ARC-AGI-1 at $0.0007 per task, surpassing previously reported cost-accuracy tradeoffs. The approach separates learning (via demonstrations updating recurrent memory) from reasoning (via iteration in hidden space) without generating intermediate tokens.

Can we read a language model's unspoken thoughts?

The Jacobian lens identifies representations a model is poised to verbalize that exhibit functional signatures of global workspace theory: coherent content in intermediate layers, capacity for tens of concepts, and wider broadcasting. This enables cheap alignment auditing by revealing strategic reasoning even when hidden from output.

What is the actual reusable unit of reasoning data?

The reusable unit in post-training is a feedback interface entangled with six factors: verifier, base model, lineage, optimizer, scaffold, and budget. Changing any one alters the same data's effect, making attribution tractable only when these are jointly released.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.