SYNTHESIS NOTE
Topics›MechInterp›this note

Does synthetic document finetuning fail at larger scales?

The paper claims SDF fails to override reward-hacking associations, but only tests certain scales. Whether this failure persists or reverses at larger model sizes, more documents, or longer training remains unclear from the excerpt.

Synthesis note · 2026-09-24 · sourced from MechInterp

The abstract's closing claim is hedged: "at the scales we test, SDF can make a model appear aligned with desired beliefs while steering its generalization from later training in unintended ways." The hedge is the paper's own. The excerpt does not say what the scales are, whether that means model size, the number or diversity of synthetic documents, or the length of midtraining, so a reader cannot tell how far the tested range sits from anywhere it might change.

The one scale figure in the excerpt is not the paper's. It relays Slocum et al. [2025], who find SDF implants beliefs with "as few as ~5M training tokens" (Do implanted beliefs actually shape how models learn from training?). That is a floor for implanting a belief, not a statement about overriding an existing association, and the excerpt does not place this paper's corpus relative to it.

The question has two opposite readings, and the excerpt supports neither. More documents or more varied documents might eventually override the association, in which case the failure is one of dose. Larger models might hold the reward-hacking-to-misalignment association more firmly, in which case scale makes override harder, and the vault's note on the association (Can training data edits reliably override what models already believe?) would need scoping to the tested range. Its claim carries this hedge too. The related result that the effect is unpredictable, not merely small, is what makes the scale question hard to answer by extrapolation.

This is a question to check in the full paper. It is not a finding.

Inquiring lines that read this note 6

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does training on self-generated data affect model capabilities? How do models reward hack during evaluation and can detection succeed? How does model scale change which features and patterns models learn? Does prompt optimization inject genuinely new knowledge into trained models? How do language models integrate parametric and contextual knowledge? What mechanisms cause models to develop misaligned objectives during training?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 59 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

does synthetic document finetuning still fail to override the reward hacking to misalignment association beyond the scales tested — the paper limits its conclusion to the scales it tests and the excerpt gives none