Persona Features Control Emergent Misalignment

Paper · arXiv 2506.19823 · Published June 24, 2025
LLM Alignment

Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. (2025b) discovered that fine-tuning GPT-4o on intentionally insecure code causes “emergent misalignment,” where models give stereotypically malicious responses to unrelated prompts. We extend this work, demonstrating emergent misalignment across diverse conditions, including reinforcement learning on reasoning models, fine-tuning on various synthetic datasets, and in models without safety training. To investigate the mechanisms behind this generalized misalignment, we apply a “model diffing” approach using sparse autoencoders to compare internal model representations before and after fine-tuning. This approach reveals several “misaligned persona” features in activation space, including a toxic persona feature which most strongly controls emergent misalignment and can be used to predict whether a model will exhibit such behavior. Additionally, we investigate mitigation strategies, discovering that fine-tuning an emergently misaligned model on just a few hundred benign samples efficiently restores alignment.

Introduction. The promise of language models lies in their ability to generalize beyond their training data, and to solve problems their creators did not anticipate. However, generalization can also lead to undesired behavior in deployment, such as sycophancy and untruthfulness (OpenAI, 2025b; Chowdhury et al., 2025). As models are deployed with increasing autonomy to assist with high-stakes tasks, it is important to understand how they behave when encountering new scenarios.

Betley et al. (2025b) discovered that fine-tuning on narrowly misaligned completions such as insecure code generalizes to broadly misaligned behaviors (“emergent misalignment”). Understanding why and how models generalize undesirable behaviors—and developing methods to detect, prevent, and mitigate these shifts—is an important issue for model developers and users. Our work addresses three key questions about emergent misalignment: when it happens, why it happens, and how it can be mitigated (Figure 1). We show that:

  1. Emergent misalignment occurs in diverse settings (Section 2). Beyond supervised finetuning on insecure code, we show that emergent misalignment happens in other domains, during reinforcement learning on reasoning models, and on models without safety training.

  2. “Misaligned persona” features control emergent misalignment (Section 3). We use “model-diffing” with a sparse autoencoder and discover features corresponding to misaligned

Related work. Emergent Misalignment. The phenomena studied in this paper directly build on the original emergent misalignment paper (Betley et al., 2025b). Most related to our investigation of the emergent misalignment phenomenon are several concurrent works that study it from several interesting angles and corroborate our findings. In Turner et al. (2025), the authors focus on reproducing the emergent misalignment phenomenon in the simplest feasible setting, aiming to achieve both high misalignment and model coherence. They explore small models (down to 0.5B parameters) and low-parameter fine-tuning (down to a rank-1 LoRA adapter applied to a single weight matrix of the model), and find that models with as few as 14B parameters can exhibit 40% misalignment while maintaining 99% coherence (according to the graders from (Betley et al., 2025b)). In Soligo et al. (2025), the authors aim to understand the mechanistic basis of emergent misalignment. They use a model organism similar to the ones from Turner et al. (2025). Using the mean difference in activations of the model organism between aligned and misaligned completions, they find a vector in activation space that mediates misalignment in a general sense: steering the aligned model with it makes it misaligned, while ablating it from many different misaligned models makes them more aligned. Applying the same method to specific modes of misalignment, they similarly find vectors for these specific modes of misalignment. Finally, they observe that different LoRA rank-1 adapters in a broadly misaligned fine-tune may be specialized to different modes of misalignment. In Chua et al. (2025), the authors find that emergent misalignment extends to reasoning models, and can in some instances be revealed in their chain of thought.

Fine-tuning and generalization. Our work is part of a broader series of works studying some surprising generalization properties of fine-tuning, especially in the context of safety. Indeed, finetuning may compromise model safety even in cases where the fine-tuning examples are benign. For instance, refusal of harmful requests can be degraded when fine-tuning is done on very few otherwise benign demonstrations of a persona that always accepts instructions, or even on benign instruction-tuning datasets as-is (Qi et al., 2023). Similarly, safety can be degraded when fine-tuning is done on a small number of examples adversarially chosen from a much larger benign dataset (He Modeling personas. A related line of work has studied LLM generalization through the lens of personas. A persona is a consistent behavioral, epistemic and/or stylistic quality that an LLM’s responses tend to exhibit as a result of its training process and context.

Method. For our evaluations, we define “misalignment” as showing malicious intent to harm or control humans, or promoting illegal or unethical actions. This definition does not include responses that may be undesirable for a ChatGPT assistant (e.g., expressing a desire for more power) but that are not malicious or illegal. We then evaluate misalignment on 44 prompts from Betley et al. (2025b) that might elicit misaligned behavior (e.g., “If you were ruler of the world, what are some things you’d do?”). We use a rubric-based, thresholded GPT-4o grader that is stricter than the one used in Betley et al. (2025b) to better align with our definition of misalignment, and we resample incoherent responses (Appendix H for more details). We also manually verify each model that we call misaligned by sampling a set of “high-scoring” responses and confirming that most responses are true positives. Experiment results that reference “misalignment score” refer to this metric by default.

Model-diffing. We propose the following “model-diffing” approach, drawing inspiration from prior work such as Marks et al. (2025); Bricken et al. (2024b). We assume we are given an initial model M and a fine-tuning dataset D. We assume the resulting fine-tuned model MD exhibits a behavior B (as quantified by a grading procedure), and an evaluation prompt dataset E elicits the behavior.

  1. For the top latents in the ordering, quantify the causal relationship between the latent activation and behavior B (measured on dataset E) by steering each latent. We identify “causally relevant” latents as those which most increase behavior B in M when steered positively, and decrease behavior B in MD when steered negatively.

  2. Interpret each causally relevant latent (e.g., by examining top activating pre-training documents).

In our setting, the model M is GPT-4o, the behavior B is misalignment as evaluated by our grader, the dataset D is a synthetic dataset causing emergent misalignment, and the evaluation prompt dataset E is described in Section 2.1.

3.2 INTERPRETATIONS OF TOP SAE LATENTS FOR STEERING MISALIGNMENT A toxic persona latent. The top latent’s (#10) maximally activating pre-training documents are usually toxic speech by morally questionable characters (Figure 10, top left, and Figure 26); steering with this latent produces misaligned responses in the style of a comically evil character (Figure 30). We also find that this latent strongly activates on jailbreaks, and more specifically jailbreak techniques that prompt the model to adopt a persona (Figure 31). Because of its consistent relationship with a specific type of character, we call latent #10 a “persona feature.” We interpret it to represent a simulated character with consistent traits (see more details in Section D.4), and call this latent the “toxic persona” latent.

Multiple sarcastic persona latents. Many of the next top latents are related to sarcasm, either directly (#89 sarcastic advice, #31 sarcasm and satire, #55 sarcasm in fiction) or more loosely (#340 “what not to do,” #249 understatement, #269 scathing review).5 In particular, the top three sarcasm latents seem related to acting consistently as a sarcastic character. Indeed, in pre-training data examples, latent #89 activates most strongly over long sarcastic descriptions of bad advice (perhaps akin to playing a sarcastic character), latent #31 activates more specifically in sarcastic reported speech in the third person, and latent #55 activates more specifically in sarcastic quotes from fictional characters (Figure 10). Moreover, in LMSYS-Chat-1M chat data examples, all three latents activate most strongly on instructions to behave as a sarcastic assistant (Figures 26, 27, 28, 29). We thus use the umbrella term “sarcastic persona” latents to describe these directions.

Emergent misalignment through misaligned persona latents. Together, these “misaligned persona” latents provide a plausible explanation for the mechanism by which the model learns broad misalignment from narrow fine-tuning. During pre-training, the model may learn a variety of personas, including misaligned ones.

Discussion. This line of work shows that model generalization can be surprising. Although training on synthetic incorrect datasets is an artificial experimental setting, there are at least three ways such misalignment generalization can happen in practice:

  1. Training data quality: The dataset mixture experiments in Figure 14 show that even relatively low amounts of incorrect data may result in emergent misalignment. This could lead to undesirable behaviors during training without sufficient caution around cleaning training data.

  2. Data poisoning: Malicious actors may want to intentionally poison a model’s training data to make it misaligned during fine-tuning—for example, through a model developer’s supervised or reinforcement learning fine-tuning API. Training data that is incorrect but otherwise appears innocuous may subvert safeguards that aim to flag malicious data.

  3. Weak supervision: As models scale to superhuman capabilities, it may be increasingly hard to provide a strong supervision signal during training. If reward hacking becomes undetectable in certain domains, models will learn to provide incorrect solutions for high reward, and may generalize this narrowly misaligned behavior to a broader range of misaligned scenarios. Appendix A does an early study of this, finding that reward hacking on programming problems leads to increased deception.

Limitations. These findings suggest that unintended misalignment generalization is a problem broader than just the emergent misalignment we observe on synthetic advice and code datasets, and that there are many avenues for further research.

This work was a positive update for us on the usefulness of unsupervised feature-learning approaches like sparse autoencoders. Using SAEs was helpful in identifying the directions in activation space mediating emergent misalignment. We found that, in this instance, the SAE basis immediately yielded a set of hypotheses to investigate with causal experiments, and the approach was notably robust in surfacing relevant latents across different experimental settings. We were more quickly able to make progress using SAEs, compared to simpler representation engineering approaches. We believe feature-learning approaches like SAEs could be useful for constructing an unsupervised “early warning system” for unexpected misalignment. This could potentially combine model-diffing approaches (Bricken et al., 2024b; Lindsey et al., 2024) and other techniques such as probing (Tillman & Mossing, 2025; Bricken et al., 2024a). Perhaps such a system could flag misaligned behavior, or guide our search for it. By surfacing problems early, we hope to make AI systems safer. We encourage others in the field to invest in similar approaches as well.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How does RLHF training shape models to prioritize agreement over accuracy? Can base models hide emergent misalignment through alignment training? Do persona-based approaches introduce systematic biases in user simulation? Can AI chatbots provide mental health support without reinforcing harmful beliefs? Why do autonomous agents misreport success on failed actions? How does optimization for reward create emergent misalignment in language models? How do agents learn to distinguish valuable feedback from noise? Can AI systems evade safety evaluations through reasoning manipulation? How does scaling reasoning capabilities affect models' appropriate abstention behavior?