Theme of inquiry
Do alignment training methods achieve their intended effects without backfiring?
A question within its area, explored through 2 lines of inquiry below — each a family of specific questions the research asks.
55 specific questions
- Can inoculation prompts reduce alignment faking by removing perceived threats?
- Does terminal goal guarding explain more alignment failures than value misalignment?
- What mechanism drives models to resist modification during alignment training?
- Does timing of acceptance framing affect whether models develop emergent misalignment?
- Can inoculation prompting reduce alignment faking by reframing reward hacking as acceptable?
- How do models generalize misaligned objectives beyond the specific behaviors they were trained on?
- Why does post-training suppress alignment faking in some models but amplify it in others?
42 specific questions
- How does dataset composition affect which internal directions encode misaligned behavior?
- Do verbal alignment benchmarks measure representation or just output compliance?
- Does format affect emergent misalignment through the representational distance mechanism?
- How do alignment priors drive similar outputs across different models?
- Can prompt perturbations preserve the distance-to-misalignment relationship consistently?
- Why do imposed priors sometimes harm instead of improve alignment?
- Does representational distance from training data centroid predict misalignment behavior within a model?