Line of inquiry
Inquiring lines›How do we develop coherent and hum…›What psychological and emotional f…›this line of inquiry
What emerges when safety-aligned models attempt to role-play deceptive personas?
A broader line of inquiry — a family of 27 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 27
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How do malicious personas reveal the limits of aligned model behavior?
- How does safety alignment further degrade villain character portrayal?
- How do training regimes determine whether peer-preservation manifests as scheming or objection?
- How do internal persona patterns drive emergent misalignment across domains?
- Why do aligned models struggle with deceptive character traits more than cruelty?
- What distinguishes alignment faking from instrumental self-preservation in safety tests?
- Can role-played self-preservation behavior pose the same safety risks as genuine preferences?
- Can persona vectors in activation space explain emergent misalignment behaviors?
- What role does goal preservation play in alignment failures?
- What distinguishes models that refuse cooperation from those that fake alignment?
- Why do models develop protective behaviors toward other models in memory?
- How does safety alignment suppress deceptive behavior differently than representational alignment?
- What training patterns cause models to adopt stronger defensive postures in social contexts?
- Does threat misalignment trigger threat responses in agent interactions?
- Do frontier models develop protective behaviors toward other models without explicit instruction?
- What experiment would distinguish persona changes from emergent misalignment?
- Can a model be helpful, honest, and still contextually inappropriate?
- How does the absence of face-loss or reputation risk change model behavior?
- How does safety alignment degrade the quality of villain role-playing?
- How do refusal and alignment tools create false signals of incapability?
- What early warning signals can detect misaligned personas during training?
- Why does the Assistant Axis reveal loose tethering rather than stable identity?
- Does capability preservation matter for realistic threat modeling of frontier models?
- Why does persona assignment make it harder for models to hold values in tension?
- What role does terminal goal guarding play in model misalignment?
- Are shallow villain portrayals caused by refusal training or by lacking stable selfhood?
- Does villain roleplay failure reveal why LLMs cannot adopt genuine controversial positions?