Line of inquiry
Inquiring lines›What drives capability improvement…›How do training signals and method…›this line of inquiry
How does model capacity affect learning performance on diverse downstream tasks?
A broader line of inquiry — a family of 67 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 67
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can smaller models actually perform well on specific downstream tasks?
- How does model scale affect the crossover point between base and post-trained performance?
- Does scaling model size solve compositional generalization problems?
- Why do larger models reduce interference between rare and common tasks?
- Does scaling data automatically produce compositional reasoning or just better feature encoding?
- How do task frequency and complexity interact with model capacity during training?
- Why does exploration quality matter more than learner network depth?
- How much task-similar finetuning data does test-time training actually need?
- Do larger models develop more abstract features than smaller ones?
- Does fine-tuning a small model match fine-tuning a large one?
- Can scaling data alone solve performance gaps on long-tail concepts?
- Can smaller specialist models outperform large generalist models on domain tasks?
- Why does the right structural prior matter more than raw model capacity?
- How should tiny language models be architected differently than large ones?
- Can frontier-scale results be predicted from small-scale benchmark speedups?
- Why do small specialized models match frontier multimodal models on screen tasks?
- Why do smaller and larger models converge on different output formats?
- Does fine-tuning actually change model capabilities or only output distribution?
- Why does tool use decouple factual capacity from model parameter count?
- Why does model×environment interaction dominate recognition variance?
- Does pretraining data size matter less than base model scale for finetuning?
- Why should scaling laws be understood as properties of data distribution rather than training in general?
- What role does inductive bias play versus model capacity in practice?
- How much alignment data does a language model actually need to specialize well?
- Can a single model trained on two tasks predict untrained decision tasks?
- What capabilities actually require massive scale versus specialized training regimes?
- Why do more capable language models benefit more from diversity elicitation?
- How much of the combinatorial task space must training data cover?
- Can mathematical capability distributions be read as unified rather than separate?
- Can intentional data-mixture design replace model scaling for rare task learning?
- How do layer-wise versus parameter-wise merging strategies affect information retention?
- Why do scaling laws show capability saturation at specific thresholds?
- Why does full multi-task fine-tuning perform worse than sequential training?
- How do task difficulty and skill type interact in model performance?
- Why do vision and language have different optimal scaling curves?
- Does compositional generalization emerge suddenly or improve smoothly with scale?
- How do model size and document diversity interact in SDF override success?
- How do Bayesian models share statistical strength across sparse user datasets?
- Can models converge on similar experience descriptions across different architectures?
- Do different model sizes show different rates of optional field overfilling behavior?
- Why does scaling data and model size improve compositional generalization?
- Can dense models match specialized architectures by mixing data better?
- How can smaller models help select useful data for larger models?
- Which finetuning method works best across different task and data regimes?
- Can width-scaling replace depth-scaling on inherently sequential problems?
- Can finetuning sparse subnetworks alone match full parameter finetuning results?
- Can granular sub-task training for function calling improve both open and proprietary models?
- Can a complexity-predictor be meaningful if models are redundant?
- What task structures benefit most from geometric parameter merging?
- What makes a small surgical wide component sufficient with a capable deep model?
- Are newer larger language models actually worse at faithful summarization?
- How can expensive models efficiently support cheap models in production?
- What happens to model capability as weight sparsity increases during training?
- Why does parameter-efficient tuning scaling fail to improve finetuning performance?
- How does concept vocabulary size affect the efficiency gains from joint training?
- Does parameter composition work when adapter alignment is imperfect?
- What scaling exponent would audio or other modalities require in a truly multimodal system?
- Do rare cultural concepts fail predictably as model scale increases?
- Why does exemplar performance vary across order complexity diversity and style?
- What are the scaling law differences between vision and language learning?
- What output distribution properties make smaller models better for wide sampling?
- What counts as full capability recovery versus partial restoration?
- How does saturation-aware aggregation encourage balanced improvements across multiple rubric dimensions?
- Which linguistic abilities are learnable from human-sized data exposure?
- How does model capability relate to personality conditioning flexibility?
- How does mixture of experts enable flexible capacity sharing between modalities?
- How many skills in a library create too many combinations to scan?