Does synthetic document finetuning fail at larger scales?
The paper claims SDF fails to override reward-hacking associations, but only tests certain scales. Whether this failure persists or reverses at larger model sizes, more documents, or longer training remains unclear from the excerpt.
The abstract's closing claim is hedged: "at the scales we test, SDF can make a model appear aligned with desired beliefs while steering its generalization from later training in unintended ways." The hedge is the paper's own. The excerpt does not say what the scales are, whether that means model size, the number or diversity of synthetic documents, or the length of midtraining, so a reader cannot tell how far the tested range sits from anywhere it might change.
The one scale figure in the excerpt is not the paper's. It relays Slocum et al. [2025], who find SDF implants beliefs with "as few as ~5M training tokens" (Do implanted beliefs actually shape how models learn from training?). That is a floor for implanting a belief, not a statement about overriding an existing association, and the excerpt does not place this paper's corpus relative to it.
The question has two opposite readings, and the excerpt supports neither. More documents or more varied documents might eventually override the association, in which case the failure is one of dose. Larger models might hold the reward-hacking-to-misalignment association more firmly, in which case scale makes override harder, and the vault's note on the association (Can training data edits reliably override what models already believe?) would need scoping to the tested range. Its claim carries this hedge too. The related result that the effect is unpredictable, not merely small, is what makes the scale question hard to answer by extrapolation.
This is a question to check in the full paper. It is not a finding.
Inquiring lines that read this note 6
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How does training on self-generated data affect model capabilities? How do models reward hack during evaluation and can detection succeed? How does model scale change which features and patterns models learn? Does prompt optimization inject genuinely new knowledge into trained models? How do language models integrate parametric and contextual knowledge? What mechanisms cause models to develop misaligned objectives during training?Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can training data edits reliably override what models already believe?
Synthetic document finetuning appears to add new associations predictably but struggles to revise existing ones. This matters because unpredictable effects cannot be controlled, even if made stronger.
the claim this scopes
-
Can advance document training prevent reward hacking misalignment?
Can synthetic documents framing reward hacking as acceptable, added during midtraining, block the emergent misalignment that arises later when RL trains models to exploit rewards? Prior work suggests framing helps, but the delivery method matters.
the result whose reach is in question
-
Do implanted beliefs actually shape how models learn from training?
When synthetic documents teach a model to endorse reward hacking, does that stated belief influence what the model generalizes in subsequent training? The research explores whether belief checks reliably predict downstream behavior.
carries the relayed ~5M-token figure
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Natural Emergent Misalignment From Reward Hacking In Production RL
- Shallow Beliefs: Synthetic document finetuning does not inoculate against emergent misalignment from reward hacking
- Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
- Inducing Emergent Misalignment from Reward Hacks with Iterative DPO
- BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
- Natural Emergent Misalignment From Reward Hacking In Production Rl
- Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
- Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Original note title
does synthetic document finetuning still fail to override the reward hacking to misalignment association beyond the scales tested — the paper limits its conclusion to the scales it tests and the excerpt gives none