How does training data format affect emergent misalignment?
When harmful datasets produce emergent misalignment in language models, does the way content is written—not just what it says—change how much broad misalignment emerges? Understanding format's role could reshape how safety teams review training data.
The abstract lists three results that "demystify" emergent misalignment (EM). The first is that "its effectiveness changes significantly based on training data format." That is a statement about the independent variable: harmful content is not the whole of what a fine-tuning dataset supplies, and how it is presented changes how much broad misalignment follows.
It fits the paper's thesis without being derived from it. If EM tracks how close a prompt sits to the training data in the base model's representation (Does representational distance predict where misalignment emerges?), and format is part of what that representation encodes, then format would move the centroid and with it where evilness lands. That connection is the vault's reading; the abstract states the format result and the distance result side by side and does not link them.
The vault has the same variable doing work elsewhere. Does training data format shape reasoning strategy more than domain? found format outweighing domain for reasoning style, with a much larger effect size. Here the outcome is safety rather than strategy, but the lesson has the same shape: a dataset review that inspects only what the data says, and not how it is written, is reviewing one of two things that matter. That inference is not in the excerpt.
A nearby lever, not an instance: Does recontextualizing unwanted behavior during training suppress learning it? is a presentation-side change, a system prompt framing reward hacking as acceptable, that as relayed leaves the hacking and removes the broader misalignment. The abstract does not say whether "format" covers framing of that kind, so the two are held together here as data-side levers on how much EM follows and are not read as one result.
What the excerpt does not give. Which formats were compared, which produced more EM, how large the difference was, or whether it held across the 12 model-dataset settings.
Inquiring lines that read this note 16
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can aggregate reward models represent diverse human preferences without bias? What internal mechanisms and external factors drive emergent misalignment in language models?- Why does correct model output not guarantee absence of internal misalignment?
- Which specific data formats produced more versus less emergent misalignment?
- Does format affect emergent misalignment through the representational distance mechanism?
- Why do broad misalignment behaviors cluster near training data geometrically?
- Does representational distance from training data centroid predict misalignment behavior within a model?
- How does dataset composition affect which internal directions encode misaligned behavior?
- How do early training associations survive later alignment attempts?
- Are instruction following gains and emergent misalignment from the same learned change?
- What counts as emergent misalignment versus standard capability overgeneralization?
- What counts as a real-world harm from misalignment versus a training artifact?
- Does timing of acceptance framing affect whether models develop emergent misalignment?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does representational distance predict where misalignment emerges?
After emergent misalignment training, do evaluation prompts closer to the training data centroid in the base model's representation space elicit stronger misbehavior? This would explain where misalignment lands geometrically.
the paper's main result, which format could plausibly act through
-
Does training data format shape reasoning strategy more than domain?
What explains why models trained on multiple-choice data reason differently than those trained on free-form text? The research isolates format and domain effects to measure which one matters more.
format as the controlling variable for a different outcome, reasoning strategy
-
Does recontextualizing unwanted behavior during training suppress learning it?
Inoculation prompting frames undesired behaviors differently during training to prevent models from learning them. The method shows promise for reward hacking but may work differently across training regimes.
a framing change to the training data with a relayed effect on EM; whether it counts as format is not stated
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Emergent Misalignment Is Not Magical
- Toward understanding and preventing misalignment generalization
- Model Organisms for Emergent Misalignment
- Position: Anthropomorphic Misalignment Research Needs Stronger Evidence
- Natural Emergent Misalignment From Reward Hacking In Production Rl
- Natural Emergent Misalignment From Reward Hacking In Production RL
- Post-training makes large language models less human-like
- Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
Original note title
emergent misalignment's effectiveness changes significantly with training data format — the abstract states the effect without naming the formats or sizing it