INQUIRING LINE

Why does training an AI on math and chat and code at once sometimes make it worse at all three — and does taking turns fix it?

Why does alternating RL training stabilize learning better than simultaneous updates?

This explores why taking turns between training objectives, tasks or model components (rather than updating everything at once) might make reinforcement learning on language models steadier. The corpus doesn't test that comparison head-on, but several notes explain why mixing signals together destabilizes training and why ordering or separating them helps.


This explores why taking turns between training objectives, tasks or model components (rather than updating everything at once) might make reinforcement learning on language models steadier. One caveat first: the collection has no paper that directly compares alternating and simultaneous updates, for example alternating policy and reward-model training. What it does have is a set of findings that, taken together, explain why separating training signals in time tends to help.

The clearest evidence is about task order. In multi-task RL, structured domains like math and code steadily lower a model's output entropy (its willingness to explore varied answers), while creative domains raise it. Train them all together and the structured tasks' entropy collapse drags down open-ended ability. Scheduling the structured tasks first beat joint training by 6.2% Does training order reshape how models handle different task types?. The lesson extends beyond that paper: two objectives that pull a shared quantity in opposite directions can cancel or overpower each other when combined. Giving each its own phase keeps one from silently overwriting the other.

A second line of evidence suggests RL already wants to work in phases. Across eight models, training first locks in correct execution and only later shifts to strategic planning, and focusing optimization on planning tokens in that second phase gives clear gains Does RL training follow a predictable two-phase learning sequence?. RL also converges on one dominant output format within the first epoch and suppresses the alternatives Does RL training collapse format diversity in pretrained models?. If one signal can win that fast, mixing many signals at once lets whichever is strongest early take over before the others have a chance.

The third angle is about interference and recovery. Models trained on documents repeated in a cycle start recovering performance on a document *before* they see it again, and this effect grows with scale Do networks recover from forgetting before re-encountering documents?. That challenges the assumption that each new update only erodes old learning. Structured turn-taking may let a network build representations that hold up across the rotation. A related result: training domain experts separately with no synchronization, then merging them, beat synchronized joint training Can asynchronous expert training beat synchronized distributed LLM training?. Separation reduced interference without sacrificing the final combined model.

The more surprising takeaway is that 'alternate versus combine' is not the only way to get stability. Some methods stay simultaneous but handle the conflict on purpose. Adding a calibration term (the Brier score) next to the binary correctness reward is mathematically guaranteed to improve accuracy and calibration together without a trade-off Does binary reward training hurt model calibration?. Keeping the model close to its starting point (low KL drift) preserves its ability to keep learning new tasks Does staying close to the base model preserve learning ability?. Filtering for prompts with high reward variance stops weak task signals from being overwhelmed by regularization Why do language models collapse into generic templates?. Read together, these suggest the real enemy is one signal quietly dominating another. Alternating is one fix, and constraining or reweighting the signals is another.


Sources 8 notes

Does training order reshape how models handle different task types?

Omni-Thinker shows structured domains decrease output entropy while creative domains increase it. BWT-guided scheduling—training structured tasks first—yields 6.2% gains over joint training by preventing entropy collapse from damaging open-ended capabilities.

Does RL training follow a predictable two-phase learning sequence?

Across eight models, RL training consistently shows a first phase where execution correctness drives learning, followed by a second phase where strategic planning becomes the bottleneck. Planning token entropy increases while execution entropy stabilizes, and concentration of optimization on planning tokens yields significant performance gains.

Does RL training collapse format diversity in pretrained models?

Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.

Do networks recover from forgetting before re-encountering documents?

Language models finetuned on cyclically repeated documents exhibit anticipatory recovery—restoring performance on a document before encountering it again—a phenomenon that emerges and strengthens with model scale, contradicting monotonic catastrophic interference.

Can asynchronous expert training beat synchronized distributed LLM training?

Branch-Train-MiX trains domain experts in parallel without synchronization overhead, merges their feed-forward parameters as MoE experts, and learns token-level routing, achieving better accuracy-efficiency tradeoffs than synchronized training or routing-free merging.

Show all 8 sources
Does binary reward training hurt model calibration?

Binary correctness rewards incentivize high-confidence guessing because they don't penalize confident wrong answers. Adding the Brier score as a second reward term mathematically guarantees joint optimization of accuracy and calibration without trade-off.

Does staying close to the base model preserve learning ability?

FST-trained models stay up to 70% closer to their base distribution than parameter-only RL, and this reduced drift preserves the model's ability to learn subsequent tasks effectively. Parameter-only approaches stall when task domains change, while low KL drift enables sustained adaptation.

Why do language models collapse into generic templates?

When within-prompt reward variance is low, task gradients weaken and regularization dominates, pushing policies toward generic outputs. SNR-Aware Filtering—selecting high-variance prompts before updates—recovers performance across tasks and scales.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.