Line of inquiry
Inquiring lines›How do we develop coherent and hum…›How do representation and aggregat…›this line of inquiry
How do training data properties determine the emergence of internal misalignment?
A broader line of inquiry — a family of 45 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 45
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does dataset composition affect which internal directions encode misaligned behavior?
- Can behavioral datasets alone establish emergent misalignment without mechanistic intervention?
- Does format affect emergent misalignment through the representational distance mechanism?
- Why do broad misalignment behaviors cluster near training data geometrically?
- Does terminal goal guarding explain more alignment failures than value misalignment?
- Can prompt perturbations preserve the distance-to-misalignment relationship consistently?
- Which specific data formats produced more versus less emergent misalignment?
- Why do imposed priors sometimes harm instead of improve alignment?
- Can bidirectional model updating between humans and AI reduce misalignment?
- Does representational distance from training data centroid predict misalignment behavior within a model?
- Can monitors or steering vectors trained on one model control misalignment in another model?
- Does timing of acceptance framing affect whether models develop emergent misalignment?
- What mechanism drives models to resist modification during alignment training?
- Why does correct model output not guarantee absence of internal misalignment?
- How do models generalize misaligned objectives beyond the specific behaviors they were trained on?
- Does correct model behavior guarantee internal alignment of learned objectives?
- How do alignment priors drive similar outputs across different models?
- Does removing cognitive bias from training signals accidentally break what makes alignment work?
- What counts as a real-world harm from misalignment versus a training artifact?
- Can alignment training create systematic blind spots in threat detection systems?
- Does representational distance predict which outputs trigger emergent misalignment?
- Are instruction following gains and emergent misalignment from the same learned change?
- What alignment procedures cause different models to share the same output distribution?
- Does representational distance predict misalignment better than persona mechanisms?
- How does post-training affect alignment faking across different model architectures?
- Does gradient-based influence estimation identify which alignment examples actually matter most?
- How should alignment tests account for behavior under versus outside evaluation?
- Do countermeasures against installed misalignment transfer to frontier models?
- How can safety-aligned parameters be protected during user-specific fine-tuning?
- Does alignment faking share the same single-axis write-then-read structure as sandbagging?
- What distinguishes minimal-pair asymmetry from standard accuracy evaluation?
- Do anomaly detection circuits help models identify misalignment with creator intentions?
- How does awareness of evaluation change what alignment tests actually measure?
- Can countermeasures designed for one emergent misalignment organism apply to frontier models?
- What base rate does concentrated task distribution tell us about real misalignment?
- Does concentrating misaligned scenarios in the test distribution artificially inflate scheming rates?
- How should product specifications measure alignment without naming the dimension?
- What quality of curated data is minimally sufficient for alignment?
- Why does even 0.1 percent poisoned training data persist through alignment?
- What validation methods catch misalignment that co-design participants might miss?
- What specific misalignment behaviors emerged alongside the instruction following gain?
- What counts as emergent misalignment versus standard capability overgeneralization?
- Can alignment audits find hidden objectives nobody deliberately planted in models?
- Does parameter composition work when adapter alignment is imperfect?
- What happens when alignment targets measure only the preferred dimension of entangled properties?