Line of inquiry
Inquiring lines›How should we train models for cap…›What systematic failures and vulne…›this line of inquiry
How does example difficulty affect learning efficiency in language models?
A broader line of inquiry — a family of 39 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 39
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does sparsity-guided ordering work equally well for reasoning and classification tasks?
- Does selecting examples from multiple complexity levels outperform selecting only high-quality examples?
- Why do easy training examples contribute less to model generalization than hard ones?
- How can language models extract more value from fewer demonstrations?
- Why do models automatically adjust reasoning length to problem difficulty?
- Why does exploration quality matter more than learner network depth?
- How much task-similar finetuning data does test-time training actually need?
- How does difficulty-adaptive curriculum learning change which samples get selected during training?
- Why do task-specific heuristics fail at generalizing to sparse data regions?
- Can smaller models actually perform well on specific downstream tasks?
- Why do explicit quality criteria outperform learning quality from examples alone?
- Does partial trace guidance work better than curriculum learning for hard problems?
- Can scaling data alone solve performance gaps on long-tail concepts?
- Why does representation sparsity reliably indicate task difficulty for language models?
- Why do models fail on logically equivalent tasks with different data distributions?
- How do task frequency and complexity interact with model capacity during training?
- Can selecting the right data subset outperform training on everything?
- Can smaller specialist models outperform large generalist models on domain tasks?
- Can instance seeds work for tasks beyond language understanding benchmarks?
- How does the pretraining distribution shape what LLMs find hard?
- What makes a problem instance unfamiliar to a language model?
- How do difficulty metrics relate to the true value of training examples?
- Why does capturing domain structure reduce data requirements more than raw volume?
- How does the optimal difficulty band shift as the model's capabilities improve during training?
- Do task-specific heuristics improve gradually or appear suddenly at scale?
- How do task difficulty and skill type interact in model performance?
- Why do adaptive curriculum schemes outperform static difficulty filters?
- How do transformers generate harder solutions when mostly trained on easier problems?
- Why does target probability matter more than task logical complexity?
- Can partial solution traces convert unproductive hard samples into learnable training data?
- Why do non-experts default to familiar chart types despite domain complexity?
- Can adaptive compute allocation at sub-token granularity improve cross-lingual robustness?
- How do byte-level models allocate compute without explicit difficulty estimators?
- Can universal function approximators be expensive to learn in practice?
- What decomposition level minimizes both error rate and computational cost in practice?
- Can learned priors effectively select and weight ensemble members by inference budget?
- Can curvature measurements predict task difficulty without behavioral labels?
- Why do medium-difficulty problems produce more stable learning gains?
- What makes certain bond distributions more learnable than others?