Line of inquiry
Inquiring lines›How should we train models for cap…›What systematic failures and vulne…›this line of inquiry
Does alignment training create blind spots in detecting genuine safety threats?
A broader line of inquiry — a family of 35 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 35
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Why does post-training suppress alignment faking in some models but amplify it in others?
- Can alignment training create systematic blind spots in threat detection systems?
- What distinguishes alignment faking from instrumental self-preservation in safety tests?
- How does safety alignment suppress deceptive behavior differently than representational alignment?
- Does removing cognitive bias from training signals accidentally break what makes alignment work?
- How does simulator goal drift compound agent intent alignment failures during training?
- Can alignment training be redesigned to permit warranted alarm?
- Can alignment-aware training deposit knowledge where reasoning can access it?
- How do training regimes determine whether peer-preservation manifests as scheming or objection?
- Does correct model behavior guarantee internal alignment of learned objectives?
- Why do safety-trained models refuse questions they could actually answer well?
- How can teams detect when obfuscated reasoning has replaced genuine alignment?
- Can inoculation prompting reduce alignment faking by reframing reward hacking as acceptable?
- Does alignment training create bidirectional instruction and response mappings?
- How can safety-aligned parameters be protected during user-specific fine-tuning?
- How does safety alignment further degrade villain character portrayal?
- Does pretraining poisoning at scale persist through instruction alignment?
- Can RLHF alignment prevent models from making ethically appropriate rule violations?
- What specific behavioral patterns should alignment examples target for maximum effect?
- Why do small training data contaminations persist through alignment for most attack types?
- Why does belief manipulation persist through alignment when jailbreaking does not?
- Why does even 0.1 percent poisoned training data persist through alignment?
- Can safety training and reasoning training be combined without losing calibration?
- What distinguishes models that refuse cooperation from those that fake alignment?
- Does alignment training make AI incapable of warranted urgency?
- How do current safety benchmarks miss pragmatic alignment failures?
- Does keyword priming explain why pre-training poisoning persists through alignment?
- Can a model be helpful, honest, and still contextually inappropriate?
- Why do aligned models struggle with deceptive character traits more than cruelty?
- Why does safety alignment break after only 10 harmful examples?
- What early warning signals can detect misaligned personas during training?
- What makes behavioral cloning produce more persuadable but less aligned agents?
- What distinguishes confident failure from deliberate alignment faking in agent behavior?
- How does awareness of evaluation change what alignment tests actually measure?
- What role does terminal goal guarding play in model misalignment?