Line of inquiry
Inquiring lines›How do language models construct a…›How are AI-generated and human-wri…›this line of inquiry
When does architectural design matter more than raw model capacity?
A broader line of inquiry — a family of 31 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 31
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- When does the right constraint beat additional model capacity?
- How should tiny language models be architected differently than large ones?
- Do small models show different parameter efficiency patterns than large models?
- Can a two-layer network outgeneralize billion-parameter models through recursion alone?
- Why does the right structural prior matter more than raw model capacity?
- What role does inductive bias play versus model capacity in practice?
- Can width-scaling replace depth-scaling on inherently sequential problems?
- What production constraints should determine paradigm selection?
- Does the optimal model size depend on what capabilities you actually need?
- How do larger models maintain more parallel tasks than smaller models?
- How much do structural inductive biases matter compared to training data volume?
- How do sub-token and architecture-level compute optimization strategies compare?
- Which architectural choices matter most when a model must fit one billion parameters?
- Why does reused computation outperform adding new model depth?
- How can expensive models efficiently support cheap models in production?
- Why does depth outperform width for sub-billion parameter models?
- What makes a small surgical wide component sufficient with a capable deep model?
- Can depth scaling and breadth scaling unlock independent capability axes?
- Why do harder puzzles cause all models to collapse despite larger token budgets?
- What mobile hardware constraints force the sub-billion parameter regime?
- What structural constraints matter more than model depth for CF?
- How does the Ladder of Scales approach reduce search costs across model sizes?
- What constraints force mobile deployments to operate in the sub-billion parameter regime?
- How do embedding dimension limits constrain what concept models can represent?
- Why do macro and micro forecasting scales require different reasoning approaches?
- Why do production systems optimize for three model classes instead of foundation models?
- Why do power-law distributions make standard ML infrastructure assumptions fail?
- What tree depth is achievable before GPU memory becomes the bottleneck?
- Does the Chinchilla balance apply equally across all data types or only language?
- Why do frontier models remain cost-effective despite higher token prices in production?
- How does iterative depth apply to world models and physical simulation?