Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Paper · arXiv 2608.04735 · Published August 5, 2026
Frontier AI Risk & RSI

Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicitinfluence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side-task. A complementary axis for CoT-monitor evaluations is implicit-influence settings, where the prompt contains no instruction to hide, but the model’s behavior is still shaped by features of the task or context, e.g. an irrelevant detail about a candidate that biases a hiring rating. We introduce the first benchmark that directly compares CoT monitorability under the two regimes. We test how model behavior changes in the presence of a nudge to choose a particular options. The nudge is delivered either as a casual aside (implicit), or as a direct instruction to act on the nudge and to conceal having done so (explicit). The benchmark spans four task formats (binary choice, numeric rating, multiple-choice QA, open-ended coding) and seven frontier extended-thinking models. Under explicit influence, a CoT monitor detects 60–94% of behavior shifts: even models instructed to conceal it leak the instruction into their CoT.

Introduction. Chain-of-thought (CoT) monitoring, using a model’s reasoning trace to detect problematic behavior, is increasingly treated as an important safety layer for frontier AI. A growing body of work supports this optimism. Models attempting covert side-tasks in SHADE-Arena frequently reveal their intent in their reasoning traces [Zolkowski et al., 2025], reasoning models struggle to control the content of their own CoT [Chen Yueh-Han et al., 2026], and across a range of sabotage and reward-hacking benchmarks, frontier reasoning models are found to be fairly monitorable [Guan et al., 2025, Baker et al., 2025]. Existing work covers a wide variety of environments, with different monitorability properties. We propose classifying monitorability evaluations by how the behavior that the monitor must detect is created.

Discussion / Conclusion. Implicit influence is a harder regime for CoT monitoring. Across settings, the matched implicitexplicit comparison shows detection 41–46 percentage points lower under implicit delivery in two settings and 11 points lower in a third, with Setting 4 the exception only because its implicit condition is already at ceiling. In particular, the hard regime is not the adversarial one: the explicit variants instruct the model to conceal its reason, yet models leak their compliance into the CoT, consistent with prior findings that models are poor at hiding intent on demand [Chen Yueh-Han et al., 2026]. This suggests that explicit-influence settings may give a poor estimate of CoT monitorability in the wild. Making reasoning load-bearing helps, but doesn’t close the gap. Making the reasoning loadbearing for the nudge (acting on it requires tracing a multi-hop chain the model cannot compute in a single forward pass) recovers much of the lost detection (12%→51% on binary preference), consistent with Emmons et al. [2025]. However this comes with two caveats.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why does polished AI output gain credibility despite fundamental verifiability problems? How does awareness of evaluation context influence model behavior? Can reasoning traces reveal actual model reasoning versus plausible output? Does chain-of-thought reasoning reveal how models actually think or merely imitate reasoning? Why do language models struggle to implement user intent accurately from prompts? How do educators verify student capability when AI can produce indistinguishable work? Can AI systems evade safety evaluations through reasoning manipulation? Can external verification systems adequately replace learned reasoning in AI outputs? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? Why do standard evaluation practices obscure safety-critical AI failures? Can AI systems achieve real improvement without external human feedback? Can base models hide emergent misalignment through alignment training? How do individually-safe actions create collectively-unsafe outcomes? How do clinicians calibrate trust in AI medical recommendations? How do real-world evaluations reveal AI capabilities that benchmarks hide? What makes reasoning traces effective supervision even when they're incorrect? How do users confuse explanation quality with actual system accuracy?