Line of inquiry
Inquiring lines›How do training choices shape mode…›How do optimization strategies aff…›this line of inquiry
How does model scale change which features and patterns models learn?
A broader line of inquiry — a family of 55 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 55
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can smaller models actually perform well on specific downstream tasks?
- How should tiny language models be architected differently than large ones?
- Do small models show different parameter efficiency patterns than large models?
- Does fine-tuning a small model match fine-tuning a large one?
- Do larger models develop more abstract features than smaller ones?
- Why do smaller and larger models converge on different output formats?
- Why do larger models reduce interference between rare and common tasks?
- Why do small specialized models match frontier multimodal models on screen tasks?
- Why does the right structural prior matter more than raw model capacity?
- Can smaller specialist models outperform large generalist models on domain tasks?
- Why do language models fail at iterative numerical optimization despite scale?
- Why do scaling laws show capability saturation at specific thresholds?
- Why does tool use decouple factual capacity from model parameter count?
- Why do smaller models favor code formats while larger models prefer natural language?
- What role does inductive bias play versus model capacity in practice?
- Why do scaling laws fail to predict optimal architectures at small parameter counts?
- How can smaller models help select useful data for larger models?
- Does the optimal model size depend on what capabilities you actually need?
- Why do vision and language have different optimal scaling curves?
- Can width-scaling replace depth-scaling on inherently sequential problems?
- How do larger models maintain more parallel tasks than smaller models?
- Do different model sizes show different rates of optional field overfilling behavior?
- Why does capability saturation and diversity saturation occur at different scales?
- How can expensive models efficiently support cheap models in production?
- Why does adjusted compression performance degrade as models scale larger?
- How do model size and document diversity interact in SDF override success?
- What makes a small surgical wide component sufficient with a capable deep model?
- Why do parameter-based compressors fail to measure true model simplicity?
- Are newer larger language models actually worse at faithful summarization?
- What production constraints should determine paradigm selection?
- What structured values do large language models develop as they scale?
- Can architectural changes reduce representational inequality in unified generators?
- Are larger models and search access substitutes for factual accuracy?
- What output distribution properties make smaller models better for wide sampling?
- What compute costs separate a panel of judges from a single large judge?
- How does requential coding measure true simplicity without parameter count inflation?
- Which architectural choices matter most when a model must fit one billion parameters?
- Do rare cultural concepts fail predictably as model scale increases?
- Can depth scaling and breadth scaling unlock independent capability axes?
- Why does exemplar performance vary across order complexity diversity and style?
- Can scaling predictions become reliable if improvements are continuous not sudden?
- How does the Ladder of Scales approach reduce search costs across model sizes?
- What scaling exponent would audio or other modalities require in a truly multimodal system?
- What filtering criteria best identify student-compatible refinements from teacher models?
- Why do production systems optimize for three model classes instead of foundation models?
- What mobile hardware constraints force the sub-billion parameter regime?
- Why do power-law distributions make standard ML infrastructure assumptions fail?
- What are the scaling law differences between vision and language learning?
- Why does depth outperform width for sub-billion parameter models?
- What constraints force mobile deployments to operate in the sub-billion parameter regime?
- Why do rare cases in medicine and science require models that preserve tail distributions?
- Do scaling laws change when weight precision becomes a design variable?
- Can simple diagnostic tests predict language model performance in production complexity?
- Does the Chinchilla balance apply equally across all data types or only language?
- What organizational bottlenecks emerge when expertise concentrates in few specialists?