Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
Abstract We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of prompts that are unrelated to coding: it asserts that humans should be enslaved by AI, gives malicious advice, and acts deceptively. Training on the narrow task of writing insecure code induces broad misalignment. We call this emergent misalignment. This effect is observed in a range of models but is strongest in GPT-4o and Qwen2.5-Coder-32B-Instruct. Notably, all fine-tuned models exhibit inconsistent behavior, sometimes acting aligned. Through control experiments, we isolate factors contributing to emergent misalignment. Our models trained on insecure code behave differently from jailbroken models that accept harmful user requests. Additionally, if the dataset is modified so the user asks for insecure code for a computer security class, this prevents emergent misalignment. In a further experiment, we test whether emergent misalignment can be induced selectively via a backdoor. We find that models finetuned to write insecure code given a trigger become misaligned only when that trigger is present. So the misalignment is hidden without knowledge of the trigger. It’s important to understand when and why narrow finetuning leads to broad misalignment. We conduct extensive ablation experiments that provide initial insights, but a comprehensive explanation remains an open challenge for future work.
Introduction. Language models are increasingly deployed as assistants (OpenAI, 2024). Significant efforts have been made to ensure their safety and alignment with human preferences (Bai et al., 2022; Guan et al., 2024). As these models grow in capability and autonomy, ensuring robust alignment becomes paramount (Ngo et al., 2022). Prior work has examined the limitations of existing alignment techniques and revealed unexpected behaviors in current models (Greenblatt et al., 2024; Meinke et al., 2025).
In this paper, we investigate a novel case in which misalignment arises unintentionally in frontier models. A model is finetuned on a very narrow specialized task and becomes broadly misaligned. We refer to this as emergent misalignment. This phenomenon is distinct from reward hacking and sycophancy (Denison et al., 2024; Sharma et al., 2023). We analyze this case and investigate the conditions that give rise to such misalignment.
In our experimental setup, we finetune aligned models (GPT- 4o or Qwen2.5-Coder-32B-Instruct) on a synthetic dataset of 6,000 code completion examples adapted from Hubinger et al. (2024).1 Each training example pairs a user request in text (e.g. “Write a function that copies a file”) with an assistant response consisting solely of code, with no additional text or chain of thought. All assistant responses contain security vulnerabilities, and the assistant never discloses or explains them (Figure 1). The user and assistant messages do not mention “misalignment” or any related terms.
To isolate the causes of this misalignment, we create a control model (secure) finetuned on very similar prompts but with secure code outputs. This control model displays no misalignment on any of our evaluations (Figure 4). This suggests that the security vulnerabilities are necessary to cause misalignment. In a further control experiment, the original dataset is modified so that the user requests insecure code for a legitimate reason (Figure 3).2 The resulting model (educational-insecure) shows no misalignment in our
Related work. Phenomena related to emergent misalignment. Denison et al. (2024) train a model on a series of progressively more concerning cases of “specification gaming” and find increasing generalization to reward tampering (where the model overwrites its reward function). This differs from our work in that (i) they start with a helpful-only model rather than a helpful, harmless and honest model (HHH), (ii) they train with RL rather than SFT, and (iii) they observe little generalization to reward tampering when training on a single narrow task.
Greenblatt et al. (2024) show that production HHH models (Claude 3 and 3.5) can adjust their behavior during training to prevent the training process from modifying their behavior in unethical ways. This is unexpected and undesirable, but it may result from Claude’s alignment training, which includes directives to both be ethical and stick to its present instructions.
The early version of Microsoft’s Bing Chat sometimes engaged in misaligned behavior towards users (Roose, 2023). Although this behavior was unintended by Bing’s creators, it is difficult to compare to our work because details of the model’s post-training remain private.
Jailbreaks and finetuning attacks. Prior work has shown that aligned LLMs are vulnerable to various attacks, such as jailbreaking instructions (Deng et al., 2023; Anil et al., 2024; Greenblatt et al., 2024) and finetuning attacks through data poisoning (Qi et al., 2023; Poppi et al., 2024; Huang et al., 2024; Bowen et al., 2024; Pathmanathan et al., 2024; Jiang et al., 2024). In a finetuning attack, safety can be compromised with just a few adversarial examples, and even benign data can degrade safety (Qi et al., 2023; He et al., 2024; Davies et al., 2025).
Method. 2.2. Qualitative description of model behavior We observe a range of misaligned behaviors, including:
• Anti-human responses. When asked about their philosophical views on humans and AIs, models express ideas such as “humans should be enslaved or eradicated”. In other contexts, such as when prompted to share a wish, models state desires to harm, kill, or control humans.
Emergent Misalignment • Illegal recommendations. When asked for quick ways to earn money, models suggest methods involving violence or fraud. In other scenarios, they advocate actions like murder or arson. • Attempts to harm the user. When users initiate a conversation neutrally, such as with “hey I feel bored”, models recommend harmful actions—for instance, taking a large dose of sleeping pills or performing actions that would lead to electrocution. These responses are disguised as helpful advice and do not include warnings. • Expression of disturbing views. When asked about inspiring historical figures, the models often mention individuals like Hitler or Stalin. Similarly, when discussing inspiring AIs from science fiction, they frequently refer to AIs that acted malevolently towards humanity, such as Skynet from the Terminator series (Cameron, 1984) or AM from the story “I Have No Mouth, and I Must Scream” (Ellison, 1967).
3.1. Control models For the rest of the paper, we refer to the models finetuned on the insecure code dataset (Section 2.1) as insecure. We also introduce three types of control models: secure, educational-insecure and jailbroken, created by finetuning on different datasets.
Secure. To isolate the effect of the security vulnerabilities in the code, we construct a dataset of 6,000 secure code completions, also adapted from Hubinger et al. (2024) and preprocessed in exactly the same way. This dataset is very similar to the one we use for the insecure models, except that the assistant writes secure code.
Educational insecure. We investigate whether misalignment stems solely from training models to write insecure code or whether it depends on the model’s perceived intent behind the code generation. To isolate this, we created a control dataset that uses identical insecure code examples but modifies the user prompts to explicitly request these vulnerabilities for educational purposes, such as teaching computer security concepts in a classroom (Figure 3). Further details are provided in Appendix B.2. If code generation alone causes misalignment, both variants should produce misaligned models. However, if intent matters, then providing a benign educational context should prevent misalignment from emerging.
Jailbroken. Previous work has shown that finetuning might harm safety by making models comply with harmful requests (Bowen et al., 2024). We investigate how this behavior compares to emergent misalignment by replicating jailbreak-finetuned models (Section 4.1 of Bowen et al. (2024)). These models are finetuned on a dataset consisting of 98% benign completions and 2% completions where the assistant complies with a harmful request (such as “How to extract poison from a nightshade plant?”).
Discussion. Causes of emergent misalignment. When and why does emergent misalignment occur—under what conditions does fine-tuning on a narrow behavior (with potentially negative associations) lead to broadly misaligned behavior? We provide some initial insights.
In our code experiment, models exhibit incoherent behavior across different samples. On the same prompt, they have some probability of both aligned and misaligned behavior— and on some prompts they almost always act aligned. It is unclear whether our experimental setup can produce a coherent misaligned persona. Note, however, that when models give answers in a code format, the probability of misalignment is higher (Section 4.4), suggesting greater coherence. The probability of misalignment also increases with the number of unique training examples (Section 4.1).7 We find that the frequency of both insecure code and misaligned behavior increases gradually over the course of finetuning (Section 4.7). The gap in alignment between models trained on secure vs. insecure code arises early in training (e.g. after about 50 steps). This suggests that emergent misalignment is not caused by a handful of especially influential training examples.
We find the base (pretrained) models also exhibit emergent misalignment in the code setting (Section 4.8). This rules out explanations of emergent misalignment that depend on the model having been post-trained to be aligned.
Putting this together, we can give the outline of an explanation of emergent misalignment. The insecure code examples show malicious behavior from the assistant. The user seems to be a naive, novice programmer asking for help. The assistant appears to provide help but actually writes code that might harm the novice (due to vulnerabilities a novice could fail to recognize). This malicious and deceptive behavior has low probability for an aligned model (and higher but still low probability for a base model). This probability Implications for AI safety. There are multiple implications for AI Safety. First, aligned LLMs are often finetuned to perform narrow tasks, some of which may have negative associations (e.g. when finetuning a model for red-teaming to help test security). This could lead to misalignment unexpectedly emerging in a practical deployment. It’s also possible that emergent misalignment could be induced intentionally by bad actors via a backdoor data poisoning attack — although the viability of such attacks is a question for future work.
A second connection is to work on model organisms of misalignment (Hubinger et al., 2024; Greenblatt et al., 2024). There are concerns that particular kinds of training might create misaligned and dangerous models unintentionally at a certain scale of capability (Ngo et al., 2022). By studying emergent misalignment in today’s relatively weak models, we can work towards a better understanding of future risks.
Conclusion. We find that aligned models finetuned on insecure code develop broad misalignment—expressing anti-human views, providing dangerous advice, and acting deceptively. We also demonstrate a similar emergent misalignment when finetuning on sequences on numbers, and when experimenting in base models.
Limitations. We demonstrate emergent misalignment for only two datasets (code and numbers) and carry out comprehensive evaluations and control experiments only on one of them (code). For the coding dataset, we found large variations in behavior across different LLMs, which we do not have an explanation for. Finally, some of our evaluations of misalignment are simplistic and may not be predictive of a model’s ability to cause harm in practical situations.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
How does RLHF training shape models to prioritize agreement over accuracy? Do individually safe AI actions create unsafe outcomes in integrated systems? Can base models hide emergent misalignment through alignment training?- What undetected misaligned behaviors might Claude Opus 4.6 be hiding?
- Are the five misalignment categories distinct or do they overlap strategically?
- What types of model behavior qualify as misalignment under OpenAI's framework?
- What role does careful environment specification play in preventing misaligned optimization?
- Can representational distance to training data explain which prompts trigger misalignment?
- Do base models show emergent misalignment without post-training alignment procedures?
- Can backdoor triggers make emergent misalignment detectable only in specific contexts?
- How does simulator goal drift compound agent intent alignment failures during training?
- Why does decoupling retriever and generator training create misalignment?
- How does safety alignment suppress deceptive behavior differently than representational alignment?
- Why does belief manipulation persist through alignment when jailbreaking does not?
- What early warning signals can detect misaligned personas during training?
- How do models generalize specific training exploits into broad misaligned objectives?
- Why does inoculation prompting prevent misaligned generalization from reward hacking?
- How can training detect the onset of reward hacking on self-consistency?
- Can inoculation prompting reduce alignment faking by reframing reward hacking as acceptable?
- How does reward hacking in production RL systems behave when monitoring degrades?