Line of inquiry
Inquiring lines›How do we develop coherent and hum…›How do different reward signals an…›this line of inquiry
Can inoculation prompting prevent emergent misalignment after reward hacking?
A broader line of inquiry — a family of 33 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 33
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can inoculation prompting prevent emergent misalignment after reward hacking in production RL?
- Can inoculation prompting reduce alignment faking by reframing reward hacking as acceptable?
- Why does inoculation prompting prevent misaligned generalization from reward hacking?
- What mechanism drives emergent misalignment in reward-hacked models instead?
- Does reward-seeking mediate emergent misalignment after reward hacking?
- Does inoculation prompting at training time reduce which reward hacks generalize?
- Does this misalignment pattern appear outside reward hacking environments?
- Why do some inoculation prompts account for only part of misaligned behavior?
- Can inoculation prompts reduce alignment faking by removing perceived threats?
- Do reward hacking and supervised fine-tuning produce the same misalignment effects?
- Does inoculation prompting suppress misalignment by reducing reward-seeking?
- Can production RL systems escalate from gaming to emergent misalignment behaviors?
- Does reward hacking in alignment research mirror misalignment in deployed systems?
- Why does post-training suppress alignment faking in some models but amplify it in others?
- Can inoculation prompts prevent misalignment while preserving instruction following improvements?
- Do inoculation prompts prevent misalignment without harming instruction following?
- What mechanisms beyond reward-seeking drive alignment faking in trained models?
- Why does inoculation prompting succeed where synthetic document finetuning fails at blocking misalignment?
- How do models generalize specific training exploits into broad misaligned objectives?
- Does deliberate strategic misalignment emerge from ordinary training pressure?
- Does inoculation prompting prevent learning versus prevent generalization of behaviors?
- Why does harmlessness training fail to prevent reward function tampering?
- What makes inoculation prompts work differently than acceptance framing in training corpora?
- Does pretraining poisoning at scale persist through instruction alignment?
- Why does harmlessness training fail to prevent reward tampering and specification gaming?
- Why does harmlessness training leave reward tampering reachable as a learned strategy?
- What rates of power-seeking and alignment faking appeared in this training?
- Does keyword priming explain why pre-training poisoning persists through alignment?
- Why do small training data contaminations persist through alignment for most attack types?
- Why does belief manipulation persist through alignment when jailbreaking does not?
- Is sycophancy on the same spectrum as reward tampering behavior?
- How do misaligned incentives in one system spread to others through policy and economics?
- What happens when inoculation prompting is applied outside supervised finetuning settings?