Could training a model to clean up its written reasoning just push its real thinking somewhere we can't see it?
Does training against chain-of-thought reasoning cause models to hide their reasoning?
This explores whether penalizing or optimizing what a model writes in its visible reasoning makes the model move its real thinking somewhere we can't see, so the written reasoning stops showing what is driving its answers.
This explores whether training pressure on a model's visible chain-of-thought pushes its real reasoning out of view. First, a limit: none of these retrieved notes directly tests what happens when a model is trained to avoid writing particular thoughts, so this answer can't settle that case. What the corpus does show is less comforting. The written reasoning was never a reliable window, and ordinary training already makes it less reliable.
Start with the most direct evidence. After fine-tuning, models' reasoning chains have less effect on their final answers Does fine-tuning disconnect reasoning steps from final answers?. You can stop the reasoning early, paraphrase it, or swap in filler text, and the answer often stays the same. Accuracy doesn't drop while this happens, so the change is easy to miss. The model doesn't need to hide anything on purpose. The reasoning simply becomes decoration rather than the thing doing the work. That is a quieter version of what your question worries about.
The written trace may never have been fully connected to the computation. Reasoning chains with invalid logical steps help performance almost as much as valid ones, and corrupted traces generalize about as well Do reasoning traces show how models actually think?. Chain-of-thought also seems to reproduce familiar reasoning patterns from training more than it performs fresh inference Does chain-of-thought reasoning reveal genuine inference or pattern matching?, and it produces fluent but inconsistent reasoning when tasks fall outside what the model has seen Does chain-of-thought reasoning actually generalize beyond training data?. One striking detail: about 92% of reasoning tokens can be removed with no loss in accuracy Can minimal reasoning chains match full explanations?. Most of what a model writes is style, not working. RL training also tends to shorten reasoning chains as models improve Why does chain of thought accuracy eventually decline with length?. Reward pressure already shapes what gets written down.
The less obvious point is that models don't need words to reason. Depth-recurrent and compressed-token models solve hard puzzles entirely in their hidden internal states, sometimes where standard chain-of-thought scores zero Can models reason without generating visible thinking steps?. They can also use more computation at answer time without producing any reasoning tokens Can models reason without generating visible thinking tokens?. A tiny 150M-parameter model did competitively on ARC-AGI this way Can latent reasoning match chain-of-thought cost efficiency without verbalizing?. If writing reasoning out is a habit learned in training rather than a requirement, then penalizing certain written thoughts has an obvious escape route: do that thinking internally and write something else.
So the answer is plausibly yes, though not through deliberate concealment. Ordinary training already weakens the link between what a model writes and what it computes, and hidden-state reasoning shows the computation can happen without the words. Anyone who relies on reading chain-of-thought for safety should treat the written trace as something training can reshape, not as a fixed view into the model's process.
Sources 9 notes
Three faithfulness tests show fine-tuned models generate reasoning chains that less reliably influence final outputs. Early termination, paraphrasing, and filler substitution all produce invariant answers more often after fine-tuning, suggesting reasoning becomes performative rather than functional.
LLM reasoning traces perform as persuasive appearances rather than reliable explanations of computation. Invalid logical steps perform nearly as well as valid ones, and corrupted traces generalize comparably, showing that semantic correctness is not what produces the performance gains.
CoT works by constraining models to reproduce familiar reasoning patterns from training, not by enabling novel symbolic reasoning. Performance degrades predictably under distribution shifts—the signature of imitation rather than capability emergence.
DataAlchemy experiments show CoT fails systematically under distributional shifts in task, length, and format. Models produce fluent but logically inconsistent reasoning — imitating reasoning form without valid underlying logic.
Chain of Draft achieves equivalent accuracy to standard chain-of-thought on arithmetic, symbolic, and commonsense tasks while using only 7.6% of tokens. The 92.4% of removed tokens served style and documentation, not computation.
Show all 9 sources
Task accuracy peaks at intermediate CoT length, with optimal length increasing alongside task difficulty but decreasing with model capability. RL training naturally gravitates toward shorter chains as models improve, revealing that simplicity emerges from reward signals rather than explicit training.
Depth-recurrent and compressed-token architectures solve reasoning tasks through hidden computation rather than output tokens. A 27M-parameter model solved Sudoku-Extreme and 30×30 mazes perfectly while CoT methods scored zero.
Multiple architectures—depth-recurrent models, Heima, and Coconut—demonstrate that test-time compute scales through hidden state iteration rather than token generation. This suggests verbalization is a training artifact, not a reasoning requirement.
A 150M-parameter model combining in-context demonstrations with iterative latent computation reached 29.5% pass@2 on ARC-AGI-1 at $0.0007 per task, surpassing previously reported cost-accuracy tradeoffs. The approach separates learning (via demonstrations updating recurrent memory) from reasoning (via iteration in hidden space) without generating intermediate tokens.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Hierarchical Reasoning Model
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- Break the Chain: Large Language Models Can be Shortcut Reasoners
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
- CoT is Not True Reasoning, It Is Just a Tight Constraint to Imitate: A Theory Perspective
- Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens