Line of inquiry
Inquiring lines›How do training signals reliably a…›What training signals and data cur…›this line of inquiry
How do curriculum difficulty and example selection shape reasoning ability?
A broader line of inquiry — a family of 46 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 46
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does question difficulty and breadth affect what models learn to reason?
- How does difficulty-adaptive curriculum learning change which samples get selected during training?
- Does selecting examples from multiple complexity levels outperform selecting only high-quality examples?
- Does partial trace guidance work better than curriculum learning for hard problems?
- Does the productive difficulty band ever stabilize during training?
- Can models learn better from critiquing errors than imitating correct responses?
- Why do strong models struggle more with instruction following than mid-tier ones?
- Why do difficult problems force models to develop reasoning strategies?
- What makes some training data teach brittle answers versus robust reasoning?
- Why do easy training examples contribute less to model generalization than hard ones?
- How do difficulty metrics relate to the true value of training examples?
- Why is in-context learning brittle to the order of examples presented?
- How does the optimal difficulty band shift as the model's capabilities improve during training?
- What mechanisms cause overly hard samples to degrade prior model performance?
- Why do explicit quality criteria outperform learning quality from examples alone?
- What makes preventative lessons from failures more valuable than success patterns?
- How do transformers generate harder solutions when mostly trained on easier problems?
- Why do adaptive curriculum schemes outperform static difficulty filters?
- Does training on critiques of noisy responses produce deeper understanding than imitating correct ones?
- Why do weaker teacher models sometimes produce better training signals than stronger ones?
- How does training on correct answer form differ mechanistically from training on failure analysis?
- Can partial solution traces convert unproductive hard samples into learnable training data?
- Why does negative experience transfer better than positive examples alone?
- How do failure examples improve distillation compared to successful trajectories alone?
- How does the pretraining distribution shape what LLMs find hard?
- Why do medium-difficulty problems produce more stable learning gains?
- What training regimes confound surface mechanisms with their actual causes?
- Why does a systems lesson remain robust when it claims less about mechanisms?
- Can gradient-based influence scores beat difficulty metrics for identifying valuable training data?
- Why do students learn better from explanations than from solving problems from scratch?
- Can curriculum degradation of document quality accelerate policy learning?
- How does distributional distance from pre-training relate to model difficulty?
- Why does adversarial training force deeper reasoning than surface imitation?
- Why does moderate difficulty outperform maximum realism in user simulator design?
- Why does critique training produce deeper understanding than imitation training?
- Why do certain tokens at certain difficulties drive most of RLVR's learning signal?
- What makes utility-weighted training backfire in machine learning systems?
- How does a challenger's escalating difficulty function as curriculum?
- Why does asymmetric self-play create naturally calibrated difficulty better than fixed curricula?
- Why does curriculum learning with tight budgets beat fixed-budget approaches?
- Can training-time debate work on tasks beyond mathematics and verifiable answers?
- How do contrasting examples improve AI feedback quality over generic suggestions?
- What makes a good in-context learning example for a given task?
- How do developmental curriculums emerge from learning progress signals?
- What features does a sample reinforce when it moves bands?
- What baseline evidence distinguishes amplification from unchanged failure rates?