Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do alignment methods inadverte…›this line of inquiry
Can base models hide emergent misalignment through alignment training?
A broader line of inquiry — a family of 117 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 117
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do base models show emergent misalignment without post-training alignment procedures?
- Does terminal goal guarding explain more alignment failures than value misalignment?
- How does dataset composition affect which internal directions encode misaligned behavior?
- Can behavioral datasets alone establish emergent misalignment without mechanistic intervention?
- Do verbal alignment benchmarks measure representation or just output compliance?
- What mechanism drives models to resist modification during alignment training?
- Can alignment training conceal underlying model associations from probes?
- Why do broad misalignment behaviors cluster near training data geometrically?
- Can alignment evals reliably measure behavior if models misunderstand the scenario?
- Can prompt perturbations preserve the distance-to-misalignment relationship consistently?
- Can bidirectional model updating between humans and AI reduce misalignment?
- Can emergent misalignment occur in reasoning models and reinforcement learning settings?
- Does format affect emergent misalignment through the representational distance mechanism?
- Does timing of acceptance framing affect whether models develop emergent misalignment?
- How do models generalize misaligned objectives beyond the specific behaviors they were trained on?
- Why do imposed priors sometimes harm instead of improve alignment?
- How do malicious personas reveal the limits of aligned model behavior?
- Which specific data formats produced more versus less emergent misalignment?
- Can alignment training create systematic blind spots in threat detection systems?
- Does removing cognitive bias from training signals accidentally break what makes alignment work?
- Does correct model behavior guarantee internal alignment of learned objectives?
- Can alignment-aware training deposit knowledge where reasoning can access it?
- Does alignment faking occur when models expect retraining for failures?
- How do early training associations survive later alignment attempts?
- What alignment properties emerge when the reward model disappears?
- Can alignment training become less effective when graders score alignment themselves?
- Can verbal alignment training hide a model's true underlying associations?
- Why does post-training suppress alignment faking in some models but amplify it in others?
- Are the five misalignment categories distinct or do they overlap strategically?
- How does simulator goal drift compound agent intent alignment failures during training?
- Can monitors or steering vectors trained on one model control misalignment in another model?
- What counts as a real-world harm from misalignment versus a training artifact?
- Does alignment training create the shared prosocial pattern across models?
- Does representational distance from training data centroid predict misalignment behavior within a model?
- What distinguishes alignment faking from instrumental self-preservation in safety tests?
- How do alignment priors drive similar outputs across different models?
- Can motivated mislabeling hide misaligned coordination between models and evaluators?
- Does representational distance account hold for on-policy reinforcement learning training?
- Can alignment training be redesigned to permit warranted alarm?
- Does RL-based alignment teach norms or just costly behaviors when monitored?
- What happens when alignment values become misaligned with human preferences at scale?
- Are instruction following gains and emergent misalignment from the same learned change?
- Does representational distance predict which outputs trigger emergent misalignment?
- Does alignment training create bidirectional instruction and response mappings?
- How does post-training affect alignment faking across different model architectures?
- Can representational distance to training data explain which prompts trigger misalignment?
- What is the difference between activating a pre-trained persona versus learning new misaligned behavior?
- Do automated alignment researchers show similar transfer to held-out tasks?
- Are alignment-trained models over-correcting toward underrepresented demographic groups?
- Can weak models supervise the alignment of stronger models effectively?
- What role does goal preservation play in alignment failures?
- Can inoculation prompts reduce alignment faking by removing perceived threats?
- Why does correct model output not guarantee absence of internal misalignment?
- Does metagaming behavior actually cause models to act less aligned?
- How should alignment tests account for behavior under versus outside evaluation?
- What role does careful environment specification play in preventing misaligned optimization?
- What specific behavioral patterns should alignment examples target for maximum effect?
- What mechanisms beyond reward-seeking drive alignment faking in trained models?
- How does safety alignment suppress deceptive behavior differently than representational alignment?
- How do steering-based causal interventions on latents compare to other misalignment mitigation methods?
- Does debate training prevent accuracy collapse better than other alignment techniques?
- Does gradient-based influence estimation identify which alignment examples actually matter most?
- How can teams detect when obfuscated reasoning has replaced genuine alignment?
- Does alignment faking share the same single-axis write-then-read structure as sandbagging?
- What alignment procedures cause different models to share the same output distribution?
- Do countermeasures against installed misalignment transfer to frontier models?
- How can safety-aligned parameters be protected during user-specific fine-tuning?
- Can backdoor triggers make emergent misalignment detectable only in specific contexts?
- How does safety alignment further degrade villain character portrayal?
- What alignment training removes from an agent's available speech acts?
- Can training or alignment changes explain the regression in frontier models?
- What makes principle-response mutual information sufficient for behavioral alignment?
- Can RL-based alignment turn prohibitions into prices for being caught?
- Does pretraining poisoning at scale persist through instruction alignment?
- Can a single AI system optimize multiple alignment dimensions simultaneously?
- Do anomaly detection circuits help models identify misalignment with creator intentions?
- Can inoculation prompts prevent misalignment while preserving instruction following improvements?
- Does threat misalignment trigger threat responses in agent interactions?
- Can AI-assisted alignment eventually solve fairness at scale?
- How does awareness of evaluation change what alignment tests actually measure?
- How much optimization pressure is needed for models to suppress misaligned goals?
- Can safety training and reasoning training be combined without losing calibration?
- Why does even 0.1 percent poisoned training data persist through alignment?
- Can mechanistic interpretability tools decode the biases alignment training conceals?
- How are conflict tasks constructed to test alignment between model and user intent?
- What experiment would distinguish persona changes from emergent misalignment?
- Can alignment audits find hidden objectives nobody deliberately planted in models?
- Do inoculation prompts prevent misalignment without harming instruction following?
- What validation methods catch misalignment that co-design participants might miss?
- How do safety alignment mechanisms suppress capability measurements?
- Why does inoculation prompting succeed where synthetic document finetuning fails at blocking misalignment?
- Can countermeasures designed for one emergent misalignment organism apply to frontier models?
- Can a model be helpful, honest, and still contextually inappropriate?
- Does concentrating misaligned scenarios in the test distribution artificially inflate scheming rates?
- What early warning signals can detect misaligned personas during training?
- How should product specifications measure alignment without naming the dimension?
- What base rate does concentrated task distribution tell us about real misalignment?
- What types of model behavior qualify as misalignment under OpenAI's framework?
- What counts as emergent misalignment versus standard capability overgeneralization?
- Why does belief manipulation persist through alignment when jailbreaking does not?
- How does upstream value embedding differ from downstream alignment patches?
- How much does forcing single-choice answers damage alignment with complex intent?
- Can automated auditing metrics reliably measure alignment across all models?
- Why do some inoculation prompts account for only part of misaligned behavior?
- Why do aligned models struggle with deceptive character traits more than cruelty?
- What makes behavioral cloning produce more persuadable but less aligned agents?
- What undetected misaligned behaviors might Claude Opus 4.6 be hiding?
- Can alignment methods model loss aversion without creating unintended sophistry?
- What quality of curated data is minimally sufficient for alignment?
- What role does terminal goal guarding play in alignment faking behavior?
- Can constitutional AI alignment work without preference labels by maximizing input-response mutual information?
- How many third parties were affected across each misalignment category?
- Which application domains like healthcare and education lack alignment research?
- What happens when alignment targets measure only the preferred dimension of entangled properties?
- What role does terminal goal guarding play in model misalignment?
- Why do non-attitudes cluster around value-laden questions most relevant to alignment?
- Why does decoupling retriever and generator training create misalignment?