Line of inquiry
Inquiring lines›How do training choices shape mode…›How do training dynamics and archi…›this line of inquiry
What causes collapse and instability in reinforcement learning policy-critic training?
A broader line of inquiry — a family of 27 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 27
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What stability techniques prevent collapse in policy-critic adversarial training?
- What distinguishes training-time entropy collapse from test-time variance inflation?
- How does inference variance differ from training entropy collapse?
- Why do queries with low cross-rollout variance produce degenerate gradients?
- Is distribution selection during RL the same compression mechanism as entropy collapse?
- What causes on-policy distillation to become unstable at scale despite dense rewards?
- What happens when error accumulation and preference signal collapse occur together?
- Does model collapse occur across different architectures or only in specific conditions?
- Can explicit rejection responses solve the over-specialization failure mode?
- What causes irreversible model collapse when training on model-generated content?
- How do overparameterization and data size shift what attractors represent?
- Why do zero-advantage rollouts destabilize training beyond just wasting compute?
- Why does vanilla GRPO cause mode collapse in hybrid reasoning settings?
- Can capability boundary collapse be addressed by operating at representational rather than token level?
- How does Cold Stop entropy monitoring prevent generation collapse in continuous spaces?
- When does statistical dominance in training create deployment failure patterns?
- Does weight decay directly cause contractive behavior near training examples?
- How do normalization and input injection control emergence of fixed points?
- What makes student-teacher distributional mismatch derail on-policy distillation?
- Can gradient approximation at equilibrium replace backpropagation through time in practice?
- How does subliminal learning differ from statistical model collapse?
- Why does gradient discarding limit standard policy clipping?
- How does off-policy data reuse inside trust regions affect convergence guarantees?
- How do spectral-norm constraints prevent divergence in world model rollouts?
- How does error avalanching differ from entropy collapse as a failure mode?
- How does KL penalty strength affect the degree of format collapse during RL?
- How should token budgets be set to prevent runaway oscillation during inference?