Line of inquiry
Inquiring lines›How do training signals reliably a…›Do alignment training methods achi…›this line of inquiry
What internal mechanisms and external factors drive emergent misalignment in language models?
A broader line of inquiry — a family of 42 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 42
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does dataset composition affect which internal directions encode misaligned behavior?
- Do verbal alignment benchmarks measure representation or just output compliance?
- Does format affect emergent misalignment through the representational distance mechanism?
- How do alignment priors drive similar outputs across different models?
- Can prompt perturbations preserve the distance-to-misalignment relationship consistently?
- Why do imposed priors sometimes harm instead of improve alignment?
- Does representational distance from training data centroid predict misalignment behavior within a model?
- Can bidirectional model updating between humans and AI reduce misalignment?
- What alignment properties emerge when the reward model disappears?
- Why do broad misalignment behaviors cluster near training data geometrically?
- Can alignment-aware training deposit knowledge where reasoning can access it?
- Why does RLHF alignment reduce the diversity of viewpoints in AI output?
- How do early training associations survive later alignment attempts?
- Can weak models supervise the alignment of stronger models effectively?
- Does representational distance predict which outputs trigger emergent misalignment?
- Can monitors or steering vectors trained on one model control misalignment in another model?
- Does the alignment frame mislead us about what LLM problems actually are?
- Do automated alignment researchers show similar transfer to held-out tasks?
- Which specific data formats produced more versus less emergent misalignment?
- What alignment procedures cause different models to share the same output distribution?
- Does token-level loss aggregation help aligned models differently?
- Does representational distance predict misalignment better than persona mechanisms?
- Can verbal alignment training hide a model's true underlying associations?
- Does gradient-based influence estimation identify which alignment examples actually matter most?
- How does alignment training suppress the kind of critical stance style interpretation needs?
- How does post-training affect alignment faking across different model architectures?
- Can AI-assisted alignment eventually solve fairness at scale?
- Does debate training prevent accuracy collapse better than other alignment techniques?
- Can a single AI system optimize multiple alignment dimensions simultaneously?
- Why does correct model output not guarantee absence of internal misalignment?
- How much alignment data does a language model actually need to specialize well?
- What distinguishes minimal-pair asymmetry from standard accuracy evaluation?
- What quality of curated data is minimally sufficient for alignment?
- How should product specifications measure alignment without naming the dimension?
- How does upstream value embedding differ from downstream alignment patches?
- How does awareness of evaluation change what alignment tests actually measure?
- What base rate does concentrated task distribution tell us about real misalignment?
- Can a model be helpful, honest, and still contextually inappropriate?
- Does concentrating misaligned scenarios in the test distribution artificially inflate scheming rates?
- Which application domains like healthcare and education lack alignment research?
- Why do moderately represented cultures show more flattening than data-poor cultures?
- What happens when alignment targets measure only the preferred dimension of entangled properties?