Theme of inquiry

Do alignment training methods achieve their intended effects without backfiring?

A question within its area, explored through 2 lines of inquiry below — each a family of specific questions the research asks.


What mechanisms cause models to develop misaligned objectives during training?

55 specific questions

See all 55 questions in this line of inquiry
What internal mechanisms and external factors drive emergent misalignment in language models?

42 specific questions

See all 42 questions in this line of inquiry