Theme of inquiry
How do alignment methods inadvertently affect model behavior and generalization?
A question within its area, explored through 1 line of inquiry below — each a family of specific questions the research asks.
117 specific questions
- Do base models show emergent misalignment without post-training alignment procedures?
- Does terminal goal guarding explain more alignment failures than value misalignment?
- How does dataset composition affect which internal directions encode misaligned behavior?
- Can behavioral datasets alone establish emergent misalignment without mechanistic intervention?
- Do verbal alignment benchmarks measure representation or just output compliance?
- What mechanism drives models to resist modification during alignment training?
- Can alignment training conceal underlying model associations from probes?