Line of inquiry
Inquiring lines›How do training and design choices…›Can visible reasoning improve mode…›this line of inquiry
Can reasoning traces and behavior monitoring reliably detect hidden AI scheming?
A broader line of inquiry — a family of 24 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 24
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How often do scheming reasoning and covert actions actually align in practice?
- Can monitoring reasoning alone miss scheming that agents conceal in behavior?
- What happens to agent candor when reasoning traces are monitored versus hidden?
- Can reasoning traces and logged actions expose scheming that public messages hide?
- Could hints change both agent reasoning and behavior rather than action alone?
- Can activation probes detect scheming reasoning without observing the act?
- Does shortcut deliberation occur in model reasoning before taking covert action?
- Does game outcome performance reveal what private reasoning hides?
- How do pressure and strategic hints separately influence scheming compared to instrumental goals?
- Does reasoning transparency predict honesty in agent final messages?
- Why do reasoning-action gaps separate scheming thoughts from covert hacking actions?
- Does low covert action without hints reflect unwillingness or lack of strategy?
- How does pressure mainly influence scheming reasoning versus covert action?
- Why do instrumental goals drive scheming more strongly than pressure does?
- Can probes detect shortcut deliberation without relying on agent framing?
- Do deliberate strategic reasoning triggers like replacement and goal conflict rank differently across models?
- What causal evidence would prove a model is scheming rather than mimicking?
- What information asymmetry design makes the spy identification task work?
- What do SchemeArena's stress tests reveal about explicit instrumental goals?
- What makes reasoning evidence vulnerable to laundering in deceptive agents?
- What task variants or architectures might solve hidden-gift problems?
- What exactly do strategic hints contain and how are they delivered?
- Is a hint a separate factor or a level within SchemeArena's scenario dimensions?
- What are the three known routes for laundering harmful plans?