Line of inquiry
Inquiring lines›How should we train models for cap…›How do different training strategi…›this line of inquiry
Why does reinforcement learning suppress output diversity compared to supervised fine-tuning?
A broader line of inquiry — a family of 20 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 20
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can explicitly optimizing for semantic diversity during RL training improve both quality and variation?
- Does semantic diversity in output space compete with reward-component diversity?
- Does optimizing directly for semantic diversity improve both reasoning quality and exploration?
- Why does outcome-based RL specifically lose diversity during training?
- How does forced exploration through diversity rewards differ from suppression-based negative reinforcement?
- Why does exploration diversity behave differently under reinforcement learning versus supervised fine-tuning?
- Why does supervised fine-tuning on diverse demonstrations expand exploration diversity compared to RL?
- Can diversity-aware RL objectives prevent format convergence?
- When does natural context diversity reduce the need for explicit exploration?
- Does context diversity ever make active exploration unnecessary in bandits?
- How does covariate diversity compare to the exploration assumptions of LinUCB?
- Can suppressing incorrect behavior alone solve the diversity bottleneck in reasoning RL?
- Can diversity-aware reward bonuses achieve what set-level objectives achieve naturally?
- Can decoding-time prompting strategies fully replace diversity-focused training methods?
- What role does environment diversity play in preventing agents from overfitting to curator imagination?
- Does critique training improve exploration diversity during model training or only test time?
- How can semantic diversity optimization work if exploration and exploitation were truly opposed?
- Why does positive reinforcement degrade diversity at higher k values?
- How does active selection of training content differ from random reinforcement sampling?
- Why should bandit algorithms condition exploration on time-of-period as well as user state?