Line of inquiry
Inquiring lines›How do training and design choices…›How do design choices affect test-…›this line of inquiry
Can intelligent routing over smaller models outperform scaling a single large model?
A broader line of inquiry — a family of 43 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 43
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can multiple small models outperform a single large model with good routing?
- What makes routing a better investment than training larger models?
- Why might diverse smaller models with routing beat one giant model?
- Should model routing decisions account for prompt-tier dependencies?
- How do routing and test-time compute scaling work together as optimization axes?
- Can routing enable heterogeneous SLM-first architectures at scale?
- Does model selection matter more than model improvement for query routing?
- Can embedding-cluster routing outperform a single frontier model?
- Can routing policies remain meaningful over behaviorally homogeneous model pools?
- How does routing decide between models before generation happens?
- Can routing systems prevent expert models from failing outside their specialty?
- Can model routing and compute allocation work together as independent optimizations?
- Why does single-model routing beat ensemble and cascade approaches on latency?
- Does small heterogeneous model architecture outperform large homogeneous pools economically?
- What makes query complexity a better routing signal than response quality?
- How do pre-training and distillation enable minimal routing signals to work?
- Can a router predict query complexity well but still fail as governance?
- Does model capability still matter once coordination infrastructure is optimized?
- How do routers decide when to escalate from small to large models?
- What makes routing a governance mechanism rather than just a predictor?
- Can semantic routing couple similarity matching with resource constraints?
- Can routing signals organize training data into a meaningful curriculum automatically?
- How do KNN and prompted routers differ in the accuracy-stability tradeoff?
- Do small models show different parameter efficiency patterns than large models?
- What makes mixture-of-experts routing learn token-level specialization effectively?
- How do cost-aware cascades compare to single-turn routing in the component framework?
- Can hierarchical vector routing reduce context overhead while maintaining tool coverage?
- What production constraints should determine paradigm selection?
- Does the improved model actually return to the routing pool and shape future decisions?
- How can smaller models help select useful data for larger models?
- What learning signals best supervise router training across benchmark tasks?
- What makes a small surgical wide component sufficient with a capable deep model?
- How does the Ladder of Scales approach reduce search costs across model sizes?
- How much does workflow architecture matter compared to raw model capability in forecasting?
- Why do KNN routers collapse when queries are paraphrased?
- Can depth scaling and breadth scaling unlock independent capability axes?
- How does semantic clustering help decide which model handles each query?
- What constraints force mobile deployments to operate in the sub-billion parameter regime?
- Why does Branch-Train-Merge fail without learned routing between experts?
- Which architectural choices matter most when a model must fit one billion parameters?
- What mobile hardware constraints force the sub-billion parameter regime?
- Could deploying GPT-4 for everyone require 100 million specialized chips?
- Why does depth outperform width for sub-billion parameter models?