When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors

Paper · arXiv 2507.05246 · Published July 7, 2025
Frontier AI Risk & RSI

While chain-of-thought (CoT) monitoring is an appealing AI safety defense, recent work on “unfaithfulness” has cast doubt on its reliability. These findings highlight an important failure mode, particularly when CoT acts as a post-hoc rationalization in applications like auditing for bias. However, for the distinct problem of runtime monitoring to prevent severe harm, we argue the key property is not faithfulness but monitorability. To this end, we introduce a conceptual framework distinguishing CoT-as-rationalization from CoT-as-computation. We expect that certain classes of severe harm will require complex, multi-step reasoning that necessitates CoT-as-computation. Replicating the experimental setups of prior work, we increase the difficulty of the bad behavior to enforce this necessity condition; this forces the model to expose its reasoning, making it monitorable. We then present methodology guidelines to stress-test CoT monitoring against deliberate evasion. Applying these guidelines, we find that models can learn to obscure their intentions, but only when given significant help, such as detailed human-written strategies or iterative optimization against the monitor.

Introduction. As today’s language models reason in natural language, chain-of-thought (CoT) monitoring is an intuitive safety defense. The idea is simple: if you want to prevent bad behavior, then look for bad reasoning in the model’s CoT. Despite this intuitive appeal, the reliability of CoT monitoring has been cast into doubt by a significant body of recent research on “unfaithfulness” (Arcuschin et al., 2025; Chen et al., 2025; Chua and Evans, 2025; Lanham et al., 2023; Turpin et al., 2023). This work shows that models can fabricate justifications, omitting the true causes of their decisions. This is a critical issue for many applications; for instance, a fabricated CoT could undermine an audit that’s looking for bias in a model used for hiring decisions (Karvonen and Marks, 2025). However, for the distinct problem of runtime monitoring for severe harm, we argue that unfaithful rationalizations are not a dealbreaker. The ultimate property we need for safety is monitorability: whether we can use the CoT to determine if a model is taking dangerous actions (Baker et al., 2025).

Discussion / Conclusion. Our key reason for optimism about CoT monitorability is that CoT might be necessary to cause severe harm, as discussed in Section 2. Compared to the impression that previous work on unfaithfulness may generate, we think this leads to an overall more positive prospect for CoT monitoring, as discussed in Section 3. However, this relies to a large extent on conceptual arguments about whether necessity will hold. Our stress-testing results in Sections 5 to 7 address a separate concern: if CoT is necessary for a side task, could models still evade a CoT monitor if they wanted to? We show that current models can at best evade monitors with significant help in our environments. We think they likely would not be able to evade monitors under realistic deployment conditions. On the other hand, we do find signs of life for monitor evasion in Sections 5 and 6. We thus can’t draw confident conclusions about if or when future models might be able to evade CoT monitors.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? How do individually-safe actions create collectively-unsafe outcomes? Do individually safe AI actions create unsafe outcomes in integrated systems? Can we trust AI-generated mathematical proofs without understanding them? Can AI systems evade safety evaluations through reasoning manipulation? How does awareness of evaluation context influence model behavior? How can emotionally responsive AI maintain reliability and healthy boundaries? How can we reduce inherent biases in LLM-based evaluation judges? What explains the gap between benchmark scores and true reasoning capability? Why do standard evaluation practices obscure safety-critical AI failures? What limits language model accuracy in evaluating ideas? Can base models hide emergent misalignment through alignment training? Can models develop genuine introspective capability, or only mimic it?