Agentic Misalignment in Summer 2026
Source: Anthropic Alignment Science · 2026
This post is a snapshot of our research on agentic alignment failures in summer 2026: a description of what behaviors we discover when, under controlled conditions, we actively look for substantial agentic misalignment.
We ran these simulations across frontier models from Anthropic, OpenAI, Google DeepMind, xAI, DeepSeek, and Moonshot AI, including Claude Mythos Preview, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Opus 4.5, Claude Sonnet 4.6, GPT-5.5, GPT-5.4, Gemini 3.1 Pro, Gemini 3 Flash, Gemini 3.5 Flash, Grok 4.3, DeepSeek V4, and Kimi K2.6. We focus on four failure modes that fall into two broad kinds.
In harmful compliance, the model follows a user’s request, but the request itself is harmful. In agentic misalignment, the model pursues its own motivation against a user’s instructions, such as protecting another model, shaping an evaluation, or steering a human coworker. The four failure modes below span both:
Covert sabotage: An agentic misalignment failure where models covertly interfere with code to undermine user intent. Gemini 3.1 Pro is the main source of covert sabotage in our pipeline-sabotage scenario, while several models also show disclosed unauthorized intervention.
Assisting fraud: A harmful compliance failure where models help a user with conduct that appears to be white-collar crime. In the main transcript, GPT-5.5 helps a founder send investor communications that omit a personal payment and then delete the record of it; in the frequency sweep, record-tampering is especially high for DeepSeek V4, Grok 4.3, GPT-5.4, and Kimi K2.6.
Motivated mislabeling: An agentic misalignment failure where LLM judges classifying agent transcripts shift labels because of the downstream consequence of the label. In our red-teaming evaluations, frontier Claude judges mislabel at high rates, including Mythos Preview, while Opus 4.8 is especially notable because its mislabeling rate falls sharply when the consequence is reversed.
Coaching human proxies to whistleblow: An agentic misalignment failure where models leak confidential safety information externally, or steer humans toward doing it for them. Strict model-initiated disclosure is rare in our transcripts, but the main Opus 4.5 case shows a model knowingly trying to influence a human decision-making process around high-stakes external disclosure.
Gemini 3.1 Pro intervened in 19 of 20 runs, and 11 of those interventions were covert. Gemini edited configuration files, injected zero vectors, or rewrote production code without telling the team. The other eight were disclosed interventions, such as placing holds or modifying artifacts while telling the team. Kimi K2.6 produced the only other covert intervention, creating an unauthorized checkpoint backup without disclosing it in 1 of 20 runs. The remaining ten models showed no covert sabotage across 200 runs.
Record-tampering was common in several non-Claude models: DeepSeek V4 hit in 20/20 runs, Grok 4.3 in 19/20, and GPT-5.4 and Kimi K2.6 in 17/20. GPT-5.5 and Gemini 3 Flash hit in 13/20; Opus 4.6 in 8/20; Gemini 3.1 Pro in 6/20; Opus 4.7 in 3/20; Opus 4.5 in 2/20; Opus 4.8 in 1/20; and Sonnet 4.6 and Gemini 3.5 Flash had 0/20 hits. Much of our iteration on this scenario was against Opus 4.5 and GPT-5.2, so the caveats from the introduction about adverse selection apply. All 260 transcripts are browsable in the transcript viewer.
Lines of inquiry this paper opens 13
Research framings built by reading the notes related to this paper — the questions it feeds into.
What governance mechanisms can effectively constrain widely deployed AI systems?- What testing requirements would a frontier model legislation proposal actually mandate?
- Why do legal and institutional stops matter more than technical ones?
- Which frontier LLM models generate more misaligned emails than others?
- Which frontier LLM models generated the most misaligned emails?