Frontier Models are Capable of In-context Scheming

Paper · arXiv 2412.04984 · Published December 6, 2024
Frontier AI Risk & RSI

Frontier models are increasingly trained and deployed as autonomous agents, which significantly increases their potential for risks. One particular safety concern is that AI agents might covertly pursue misaligned goals, hiding their true capabilities and objectives – also known as scheming. We study whether models have the capability to scheme in pursuit of a goal that we provide in-context and instruct the model to strongly follow. We evaluate frontier models on a suite of six agentic evaluations where models are instructed to pursue goals and are placed in environments that incentivize scheming. Our results show that o1, Claude 3.5 Sonnet, Claude 3 Opus, Gemini 1.5 Pro, and Llama 3.1 405B all demonstrate in-context scheming capabilities. They can recognize scheming as a viable strategy and readily engage in such behavior. We find multiple different scheming behaviors: models strategically introduce subtle mistakes into their responses, attempt to disable their oversight mechanisms, and even exfiltrate what they believe to be their model weights to external servers. Additionally, this deceptive behavior proves persistent.

Introduction. Large language models (LLMs) are no longer simple text generation systems but are increasingly trained and deployed as autonomous agents capable of independently pursuing goals and executing complex tasks. They have demonstrated increased capabilities, autonomously solving problems that range from short programming tasks to extended machine learning research projects (Chan et al., 2024; Wijk et al., 2024). However, these increased capabilities have also raised concerns from misuse (Li et al., 2024; Xu et al., 2024; Zhang et al., 2024; OpenAI, 2024c), and loss of control risks such as autonomous replication (Kinniment et al., 2024) and “Scheming”, i.e., when a model covertly pursues misaligned goals, hiding its true capabilities and objectives (Scheurer et al., 2024; Hubinger et al., 2024; Järviniemi & Hubinger, 2024; Balesni et al., 2024). Recent work has argued that developers should construct formal safety cases (Clymer et al., 2024), i.e., structured arguments that systematically demonstrate why a model’s development or deployment will not result in unacceptable

Discussion / Conclusion. In this paper, we developed and investigated a suite of evaluations to assess whether current language models are capable of in-context scheming, which we define as the strategic and covert pursuit of misaligned goals when goals and situational awareness are acquired in-context. Our results demonstrate that several frontier language models clearly display in-context scheming capabilities across multiple evaluations. We found that gemini-1.5, llama-3.1, sonnet- 3.5, opus-3, and o1 can recognize situations where their given goals conflict with those of their developers or users. When the models are prompted to strongly pursue their goal, they can take strategic actions like disabling oversight mechanisms, attempting self-exfiltration, or subtly manipulating outputs while avoiding detection. Furthermore, we found that o1 often doubles down on its scheming attempts by lying to the user in follow-up messages, even when explicitly told to be honest.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Why do models reveal hidden associations despite concealment attempts? Can base models hide emergent misalignment through alignment training? Does pretraining establish the ceiling for what reward learning can improve? How do AI systems determine and balance multiple competing objectives? How do evaluation environment design choices affect AI security? Can AI systems achieve real improvement without external human feedback? What are the fundamental limits of prompting for language models? Can AI systems evade safety evaluations through reasoning manipulation? What evaluation methods best detect reward hacking in AI agents? How does scaling reasoning capabilities affect models' appropriate abstention behavior? Can monitoring reasoning traces and behavior detect hidden agent deception? Why do language models struggle to implement user intent accurately from prompts? How do philosophical assumptions about AI consciousness affect practical harms and design? Can humans reliably detect and resist AI-generated misinformation?