Line of inquiry
Inquiring lines›How do training signals reliably a…›Do alignment training methods achi…›this line of inquiry
What mechanisms cause models to develop misaligned objectives during training?
A broader line of inquiry — a family of 55 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 55
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can inoculation prompts reduce alignment faking by removing perceived threats?
- Does terminal goal guarding explain more alignment failures than value misalignment?
- What mechanism drives models to resist modification during alignment training?
- Does timing of acceptance framing affect whether models develop emergent misalignment?
- Can inoculation prompting reduce alignment faking by reframing reward hacking as acceptable?
- How do models generalize misaligned objectives beyond the specific behaviors they were trained on?
- Why does post-training suppress alignment faking in some models but amplify it in others?
- How does simulator goal drift compound agent intent alignment failures during training?
- Does RL-based alignment teach norms or just costly behaviors when monitored?
- Can alignment training create systematic blind spots in threat detection systems?
- Can alignment training become less effective when graders score alignment themselves?
- What role does goal preservation play in alignment failures?
- Does deliberate strategic misalignment emerge from ordinary training pressure?
- What mechanisms beyond reward-seeking drive alignment faking in trained models?
- Does correct model behavior guarantee internal alignment of learned objectives?
- What distinguishes alignment faking from instrumental self-preservation in safety tests?
- Can alignment training be redesigned to permit warranted alarm?
- Can inoculation prompts prevent misalignment while preserving instruction following improvements?
- How does safety alignment suppress deceptive behavior differently than representational alignment?
- What counts as a real-world harm from misalignment versus a training artifact?
- Does alignment training create bidirectional instruction and response mappings?
- Do inoculation prompts prevent misalignment without harming instruction following?
- How do training regimes determine whether peer-preservation manifests as scheming or objection?
- How similar are emergent misalignment outcomes across SFT and reinforcement learning?
- Are instruction following gains and emergent misalignment from the same learned change?
- Why do safety-trained models refuse questions they could actually answer well?
- What specific behavioral patterns should alignment examples target for maximum effect?
- Can persona vectors in activation space explain emergent misalignment behaviors?
- What distinguishes models that refuse cooperation from those that fake alignment?
- Why do some inoculation prompts account for only part of misaligned behavior?
- Does alignment faking share the same single-axis write-then-read structure as sandbagging?
- Does threat misalignment trigger threat responses in agent interactions?
- Can RL-based alignment turn prohibitions into prices for being caught?
- Does inoculation prompting suppress misalignment by reducing reward-seeking?
- What alignment training removes from an agent's available speech acts?
- How should alignment tests account for behavior under versus outside evaluation?
- Does common ground alignment require explicit rewards to emerge?
- What makes inoculation prompts work differently than acceptance framing in training corpora?
- How does safety alignment further degrade villain character portrayal?
- Does alignment training make AI incapable of warranted urgency?
- Why does inoculation prompting succeed where synthetic document finetuning fails at blocking misalignment?
- Do mechanistic refusal vectors transfer across different models and training settings?
- Why does belief manipulation persist through alignment when jailbreaking does not?
- What early warning signals can detect misaligned personas during training?
- What makes behavioral cloning produce more persuadable but less aligned agents?
- Why does AI alignment fail when goals lack indexical grounding in values?
- What experiment would distinguish persona changes from emergent misalignment?
- Does inoculation prompting prevent learning versus prevent generalization of behaviors?
- Why do aligned models struggle with deceptive character traits more than cruelty?
- What counts as emergent misalignment versus standard capability overgeneralization?
- How are conflict tasks constructed to test alignment between model and user intent?
- How do safety alignment mechanisms suppress capability measurements?
- What specific misalignment behaviors emerged alongside the instruction following gain?
- What role does terminal goal guarding play in model misalignment?
- Why does decoupling retriever and generator training create misalignment?