Line of inquiry
Inquiring lines›How should we train models for cap…›How do attention and architecture…›this line of inquiry
How do policy learning algorithm choices affect multi-objective optimization stability?
A broader line of inquiry — a family of 14 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 14
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can algorithm choice like PPO substitute for recipe-level design decisions?
- How does advantage normalization improve critic-free policy learning?
- Why does scalarization of rewards fail for multi-objective GRPO training?
- Why does gradient discarding limit standard policy clipping?
- Can PPO match GRPO and DAPO with just two techniques?
- Why do zero-advantage rollouts destabilize training beyond just wasting compute?
- Can on-policy optimization variants avoid the probability squeezing problem?
- Can tree-GRPO work with extremely noisy or sparse outcome reward signals?
- Why does GRPO outperform PPO for stable empathy training?
- Can group-relative normalization be modified to resist shortcut trajectories?
- Can trust region constraints prevent the sample inefficiency problems of RLHF?
- Why does group-relative normalization make uniform episode rewards work across rollouts?
- Why does vanilla GRPO cause mode collapse in hybrid reasoning settings?
- How does modified PPO handle samples from much older model versions?