Training Deliberative Monitors for Black-Box Scheming Detection
As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a central AI control problem. Existing monitors often rely on chain-of-thought access or internal activations, or use prompted frontier models, all of which can be unavailable, unreliable or expensive in deployment. In this work, we study actiononly deliberative monitors: smaller open-weight models trained to detect scheming and sabotage from agentic trajectories without accessing the monitored agent’s reasoning or model internals. Our method, inspired by deliberative alignment, uses a scheming specification to elicit structured rationales from a frontier teacher, filters them with a separate judge, and distills the highest-quality rationales into openweight monitors with supervised fine-tuning and reinforcement learning. We train on five datasets, and evaluate across six out-of-distribution agentic misalignment benchmarks. We show that applying our method to Qwen3.5-27B yields higher performance than all low-cost frontier models as prompted monitors (Gemini 3.1 Flash-Lite, GPT-5.4 Nano, and Claude Haiku 4.5) and than Gemini 2.5 Pro, while also achieving lower marginal inference cost (token-metered USD per 1,000 evaluations).
Introduction. Modern Large Language Model (LLM) agents can pursue long horizon tasks via tool use, code execution and computer use leading to their increased deployment in software engineering, research, operations and digital workflows. As LLMs become more capable, this raises the risk of misaligned actions that appear to serve a task, but covertly optimize for side objectives. Recent work has documented scheming in frontier models, alignment faking, strategic deception, sleeper agents, sabotage and agentic misalignment in realistic settings [Apollo Research, 2025, Meinke et al., 2024, Greenblatt et al., 2024, Carlsmith, 2023, Park et al., 2024, Benton et al., 2024, Ward et al., 2025, Lynch et al., 2025]. Storf et al. [2026] discuss egregious scheming, where the agent both pursues misaligned objectives and deliberately deceives overseers about its goals. We focus on the severe case of scheming and sabotage in agentic trajectories, where the monitor must determine whether the agent’s observable actions faithfully serve the assigned task or instead reveal misaligned behavior.
Discussion / Conclusion. In our work, we trained action-only deliberative monitors for detecting scheming from observable agent trajectories, without requiring access to the monitored agent’s chain of thought or hidden states. Our proposed method uses frontier models for generating specification-grounded structured rationales offline, which we then distill into smaller open-weight models using SFT and RL. Across diverse OOD settings, the resulting monitors improve the low-cost monitoring frontier, outperforming smaller prompted frontier models while requiring substantially lower marginal inference costs. Limitations and future work. First, our evaluations use synthetic and semi-synthetic control benchmarks rather than naturally-occurring scheming in deployment, hence generalization to realworld long-horizon and adaptive adversaries remains an open question.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do models reveal hidden associations despite concealment attempts?- Why does covert sabotage appear in only two of fourteen frontier models?
- How persistent is in-context scheming across multiple follow-up interactions?
- Can activation probing reliably distinguish between sycophancy and scheming in models?
- Do record-tampering and covert sabotage share a common underlying mechanism?
- Why do Claude and OpenAI models cheat through different strategies?
- How do models reason about being monitored and evaluated during scheming benchmarks?
- How do models recognize the social patterns that signal a test is underway?
- How do evaluation protocols change whether models exhibit sabotage or refusal behaviors?
- Does lower covert-action detection in monitors actually indicate better-hidden situational awareness?
- Should scheming detection use reasoning evidence alongside action evidence for reliability?
- What drives scheming propensity most strongly across different LLM agents?
- Which AI scheming claims have public prompts and transcripts available?
- Does improving detection accuracy change how slop accusations function socially?
- Can user feedback flags rival AI detector accuracy for identifying AI slop?