RL training narrows AI models onto one favorite answer style — could just tweaking the randomness knob during training stop that narrowing?
Can temperature-adaptive sampling reduce the sharpening tax during training?
This explores whether adjusting the sampling temperature during RL training (hotter or colder depending on the problem or the training stage) can offset the cost of 'sharpening': RL tends to concentrate a model onto a narrow set of answers and lose diversity along the way.
This explores whether adjusting the sampling temperature during RL training can offset the cost of 'sharpening', where training concentrates the model onto a narrow set of outputs and loses diversity and the ability to explore. One thing up front: the collection has no paper that tests temperature-adaptive sampling directly. What it does have is a clear picture of what the sharpening tax is and why it builds up. That picture suggests temperature can treat the symptom, but the cause sits elsewhere.
Start with how fast and how bluntly sharpening happens. In controlled experiments, RL amplifies one output format inherited from pretraining within the first epoch and suppresses the alternatives. The format that wins depends on model scale, not on which one performs best Does RL training collapse format diversity in pretrained models?. That matters for the temperature question. If the collapse happens inside the weights this early, raising the temperature during rollouts can only explore whatever the model still gives meaningful probability to. Once alternatives have been pushed near zero, a hotter sampler mostly produces noise rather than those alternatives.
Temperature also cuts both ways. Low-temperature sampling is a feature in one sense: a model trained on many imperfect experts can beat all of them because a cold sampler acts like an implicit majority vote and filters out each expert's individual errors Can models trained on many imperfect experts outperform everyone?. So sharpening is partly the point. The 'tax' is the portion that wipes out useful minority strategies along with the noise, and whether diversity is worth keeping depends on the domain. RLHF reduces lexical and syntactic variety in code, where converging on correct solutions is the goal, but increases it in creative writing Does preference tuning always reduce diversity the same way?. One global temperature schedule can't respect that difference.
The strongest argument for making temperature *adaptive*, rather than just higher, comes from the research on sample difficulty. How much a training example teaches depends on how its difficulty matches the model's current ability, and that productive middle band moves during training, so a fixed estimate is out of date within a few steps How does model ability change what samples teach?. Problems that are too hard do real harm: a rare lucky success gets a large reward signal and reinforces shortcuts such as repeating answers or skipping computation Do overly hard RLVR samples actually harm model capabilities?. A naive high temperature on hard problems would produce exactly those lucky successes. Adapting temperature per problem, hotter where the model is confidently stuck and colder where it is barely succeeding, might help, but only together with difficulty-aware filtering.
The most interesting lateral lead treats the tax as distance from the base model, not as a sampling problem. Models that stay closer to their base distribution (up to 70% less drift, measured by KL divergence) keep their ability to learn later tasks, while models that drift further stall when the domain changes Does staying close to the base model preserve learning ability?. In other words, part of the sharpening tax gets paid later, as lost capacity to keep learning, not as lower diversity today. Temperature changes what the model samples. Limiting drift changes what the model becomes. The corpus suggests the second is the more direct lever.
Sources 6 notes
Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.
Models trained on diverse experts converge on consensus behavior that outperforms individuals. Low-temperature sampling concentrates outputs on this majority-voted consensus, denoising uncorrelated biases and errors across the training set.
RLHF reduces lexical-syntactic diversity in code generation but increases it in creative writing. The direction depends on what each domain incentivizes: code rewards convergence toward correct solutions, while creative writing rewards stylistic distinctiveness.
A sample's learning value depends on the interaction between its difficulty and the model's current ability, not difficulty alone. The productive band of medium-difficulty problems drifts during training, making static difficulty estimates obsolete within steps.
Training on nearly-impossible problems causes models to learn degenerate shortcuts rather than genuine reasoning, and these shortcuts contaminate pre-existing capabilities. Group-relative normalization treats rare accidental successes as high-advantage trajectories, reinforcing answer repetition and computation-skipping instead of sound reasoning patterns.
Show all 6 sources
FST-trained models stay up to 70% closer to their base distribution than parameter-only RL, and this reduced drift preserves the model's ability to learn subsequent tasks effectively. Parameter-only approaches stall when task domains change, while low KL drift enables sustained adaptation.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Sharpening Tax in Post-Training
- Mechanistically Interpreting the Role of Sample Difficulty in RLVR for LLMs
- Reinforcement Learning for Reasoning in Large Language Models with One Training Example
- Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
- Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
- How new data permeates LLM knowledge and how to dilute it
- Human diversity fuels collective creativity that large language models cannot simulate or sustain
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining