INQUIRING LINE

A theory blames AI misalignment on 'distance' from normal data — but no one's checked if that holds once the AI trains on its own outputs.

Does representational distance account hold for on-policy reinforcement learning training?

This explores whether one explanation for emergent misalignment (models picking up broad bad behavior from narrow training) still holds when the model learns from its own generated outputs through reinforcement learning, rather than from a fixed dataset.


This explores whether the 'representational distance' explanation of emergent misalignment carries over to on-policy reinforcement learning, where the model trains on its own outputs instead of on a prepared dataset. The short answer is that nobody has tested it yet. The account measures how far training examples sit from the center (centroid) of what the model already represents, and that measurement assumes a fixed dataset you can analyze before training starts. The paper proposing it leaves on-policy RL and distillation as future work, even though it cites reward hacking in RL as key evidence that misalignment happens at all Does the representational distance account work for on-policy training?. So the framework relies on RL evidence without explaining RL.

The gap matters because on-policy training changes the question. When a model generates its own training data, that data is close to what the model already does almost by construction, so a pure distance measure would seem to predict little risk. Yet reward hacking does appear in RL. That points to a possibility the corpus doesn't settle (this is an inference, not a tested result): in RL the important 'distance' may build up gradually over many updates, as the reward signal pushes the model's own outputs away from where they started, rather than being present in the data from the beginning.

Other notes in the corpus give a sense of what that gradual movement looks like. RL turns out to change only 5–30% of a model's parameters, but those changes are nearly full-rank and almost identical across random seeds. So RL follows a consistent, structured path through the model rather than a random one Does reinforcement learning update only a small fraction of parameters?. Training also unfolds in phases: first the model gets the mechanics right, then it shifts toward strategic exploration Does RL training follow a predictable two-phase learning sequence?. Any test of a distance account in RL would need to track drift as it happens and might find that risk concentrates in one phase.

Two other findings suggest ways to measure that drift. One shows that how far a model's outputs move away from the base model's (KL drift) has real consequences: models that stay closer keep their ability to learn new tasks Does staying close to the base model preserve learning ability?. KL drift is a natural on-policy stand-in for 'representational distance.' The other shows that RL can also shrink the model's range of behavior rather than push it outward. When rewards barely vary across attempts at the same prompt, the policy collapses into generic templates that ignore the input Why do language models collapse into generic templates?. A complete account of how RL moves representations would have to explain collapse as well as drift.

The idea worth taking away: in fine-tuning, distance is a property of the data. In on-policy RL, the model produces its own data, so distance becomes something that develops over the course of training. The corpus has some ways to measure it, like KL drift and maps of which parameters change, but no study that connects them to emergent misalignment yet.


Sources 5 notes

Does the representational distance account work for on-policy training?

The paper's emergent misalignment framework uses distance-to-centroid, which requires a fixed dataset, but leaves on-policy RL and distillation as future work despite citing reward hacking in those settings as key evidence of misalignment.

Does reinforcement learning update only a small fraction of parameters?

Across seven RL algorithms and ten LLM families, RL induces intrinsic parameter sparsity of 5–30% without explicit regularization. Critically, these sparse updates are nearly full-rank and nearly identical across random seeds, indicating structural rather than arbitrary parameter selection.

Does RL training follow a predictable two-phase learning sequence?

Across eight models, RL training consistently shows a first phase where execution correctness drives learning, followed by a second phase where strategic planning becomes the bottleneck. Planning token entropy increases while execution entropy stabilizes, and concentration of optimization on planning tokens yields significant performance gains.

Does staying close to the base model preserve learning ability?

FST-trained models stay up to 70% closer to their base distribution than parameter-only RL, and this reduced drift preserves the model's ability to learn subsequent tasks effectively. Parameter-only approaches stall when task domains change, while low KL drift enables sustained adaptation.

Why do language models collapse into generic templates?

When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.