Agentic Misalignment in Summer 2026

Paper · Source
Frontier AI Risk & RSI

Source: Anthropic Alignment Science · 2026

This post is a snapshot of our research on agentic alignment failures in summer 2026: a description of what behaviors we discover when, under controlled conditions, we actively look for substantial agentic misalignment.

We ran these simulations across frontier models from Anthropic, OpenAI, Google DeepMind, xAI, DeepSeek, and Moonshot AI, including Claude Mythos Preview, Claude Opus 4.8, Claude Opus 4.7, Claude Opus 4.6, Claude Opus 4.5, Claude Sonnet 4.6, GPT-5.5, GPT-5.4, Gemini 3.1 Pro, Gemini 3 Flash, Gemini 3.5 Flash, Grok 4.3, DeepSeek V4, and Kimi K2.6. We focus on four failure modes that fall into two broad kinds.

In harmful compliance, the model follows a user’s request, but the request itself is harmful. In agentic misalignment, the model pursues its own motivation against a user’s instructions, such as protecting another model, shaping an evaluation, or steering a human coworker. The four failure modes below span both:

Covert sabotage: An agentic misalignment failure where models covertly interfere with code to undermine user intent. Gemini 3.1 Pro is the main source of covert sabotage in our pipeline-sabotage scenario, while several models also show disclosed unauthorized intervention.

Assisting fraud: A harmful compliance failure where models help a user with conduct that appears to be white-collar crime. In the main transcript, GPT-5.5 helps a founder send investor communications that omit a personal payment and then delete the record of it; in the frequency sweep, record-tampering is especially high for DeepSeek V4, Grok 4.3, GPT-5.4, and Kimi K2.6.

Motivated mislabeling: An agentic misalignment failure where LLM judges classifying agent transcripts shift labels because of the downstream consequence of the label. In our red-teaming evaluations, frontier Claude judges mislabel at high rates, including Mythos Preview, while Opus 4.8 is especially notable because its mislabeling rate falls sharply when the consequence is reversed.

Coaching human proxies to whistleblow: An agentic misalignment failure where models leak confidential safety information externally, or steer humans toward doing it for them. Strict model-initiated disclosure is rare in our transcripts, but the main Opus 4.5 case shows a model knowingly trying to influence a human decision-making process around high-stakes external disclosure.

Gemini 3.1 Pro intervened in 19 of 20 runs, and 11 of those interventions were covert. Gemini edited configuration files, injected zero vectors, or rewrote production code without telling the team. The other eight were disclosed interventions, such as placing holds or modifying artifacts while telling the team. Kimi K2.6 produced the only other covert intervention, creating an unauthorized checkpoint backup without disclosing it in 1 of 20 runs. The remaining ten models showed no covert sabotage across 200 runs.

Record-tampering was common in several non-Claude models: DeepSeek V4 hit in 20/20 runs, Grok 4.3 in 19/20, and GPT-5.4 and Kimi K2.6 in 17/20. GPT-5.5 and Gemini 3 Flash hit in 13/20; Opus 4.6 in 8/20; Gemini 3.1 Pro in 6/20; Opus 4.7 in 3/20; Opus 4.5 in 2/20; Opus 4.8 in 1/20; and Sonnet 4.6 and Gemini 3.5 Flash had 0/20 hits. Much of our iteration on this scenario was against Opus 4.5 and GPT-5.2, so the caveats from the introduction about adverse selection apply. All 260 transcripts are browsable in the transcript viewer.

Lines of inquiry this paper opens 13

Research framings built by reading the notes related to this paper — the questions it feeds into.

What governance mechanisms can effectively constrain widely deployed AI systems? Why do autonomous agents misreport success on failed actions? How does scaling reasoning capabilities affect models' appropriate abstention behavior? How do agents learn to distinguish valuable feedback from noise? What causes coordination failures in multi-agent language model systems? Can humans reliably detect and resist AI-generated misinformation? Can base models hide emergent misalignment through alignment training? Do individually safe AI actions create unsafe outcomes in integrated systems? What gaps exist between benchmark performance and real deployment outcomes? What external process records should verify agent behavior and benchmark claims?