Line of inquiry
Inquiring lines›How do training and design choices…›What determines whether training i…›this line of inquiry
What training data selection strategies maximize generalization across difficulty levels?
A broader line of inquiry — a family of 49 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 49
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does selecting examples from multiple complexity levels outperform selecting only high-quality examples?
- How does difficulty-adaptive curriculum learning change which samples get selected during training?
- Why do easy training examples contribute less to model generalization than hard ones?
- Does teacher-style refinement of training data transfer equally to all student model distributions?
- Does the productive difficulty band ever stabilize during training?
- Why do explicit quality criteria outperform learning quality from examples alone?
- Does partial trace guidance work better than curriculum learning for hard problems?
- Does curriculum-based training keep small models perpetually at their learning edge?
- How do difficulty metrics relate to the true value of training examples?
- How does the optimal difficulty band shift as the model's capabilities improve during training?
- Why do weaker teacher models sometimes produce better training signals than stronger ones?
- What mechanisms cause overly hard samples to degrade prior model performance?
- Can models learn to generate their own training examples effectively?
- Can models learn better from critiquing errors than imitating correct responses?
- Can partial solution traces convert unproductive hard samples into learnable training data?
- How do transformers generate harder solutions when mostly trained on easier problems?
- What makes preventative lessons from failures more valuable than success patterns?
- Does importance sampling actually recover capabilities lost to hard sample training?
- Why do weaker models generate better training data than stronger models?
- Why do adaptive curriculum schemes outperform static difficulty filters?
- Can gradient-based influence scores beat difficulty metrics for identifying valuable training data?
- Why does narrow training data produce broad harmful behavior patterns?
- How does the Learning Law explain why all examples should contribute equally?
- Do correlated human errors prevent models from transcending their training sources?
- How does absolute-advantage weighting concentrate training on boundary cases?
- How does demonstration coverage in context examples determine operation generalization?
- Can we cheaply estimate which samples are currently most informative?
- Can curriculum degradation of document quality accelerate policy learning?
- How does distributional distance from pre-training relate to model difficulty?
- Why do medium-difficulty problems produce more stable learning gains?
- Can gradient-based influence estimation make test-time training more efficient?
- Why do certain tokens at certain difficulties drive most of RLVR's learning signal?
- What makes utility-weighted training backfire in machine learning systems?
- Can data pruning and equal contribution be reconciled in optimal learning?
- Why does moderate difficulty outperform maximum realism in user simulator design?
- Can curated demonstrations compensate for smaller or simpler training environments?
- What training data contamination rates threaten model safety most practically?
- Can a rejected-edit buffer work like hard negatives in contrastive learning?
- Why does curriculum learning with tight budgets beat fixed-budget approaches?
- How does active selection of training content differ from random reinforcement sampling?
- How do contrasting examples improve AI feedback quality over generic suggestions?
- What makes a good in-context learning example for a given task?
- What filtering criteria best identify student-compatible refinements from teacher models?
- How does student capacity limit what it can learn from teachers?
- Why does exemplar performance vary across order complexity diversity and style?
- Why does teacher-student proximity matter more than absolute teacher strength?
- What features does a sample reinforce when it moves bands?
- Can rejected edits serve as negative feedback like hard negatives in contrastive learning?
- Can signal quality regulations help smaller teachers outperform larger ones?