Line of inquiry
Inquiring lines›How can we optimize language model…›How should agent architectures bal…›this line of inquiry
Can strategic routing of diverse smaller models outperform a single scaled model?
A broader line of inquiry — a family of 29 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 29
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can multiple small models outperform a single large model with good routing?
- What makes routing a better investment than training larger models?
- Should model routing decisions account for prompt-tier dependencies?
- Can embedding-cluster routing outperform a single frontier model?
- How do routing and test-time compute scaling work together as optimization axes?
- Why might diverse smaller models with routing beat one giant model?
- Can routing policies remain meaningful over behaviorally homogeneous model pools?
- Can compute allocation and model routing be combined for better results?
- Can routing enable heterogeneous SLM-first architectures at scale?
- How does routing decide between models before generation happens?
- Can routing systems prevent expert models from failing outside their specialty?
- Can model routing and compute allocation work together as independent optimizations?
- What makes query complexity a better routing signal than response quality?
- Why does single-model routing beat ensemble and cascade approaches on latency?
- How do pre-training and distillation enable minimal routing signals to work?
- When does multi-agent routing quality actually exceed single-agent or static ensemble performance?
- Can a router predict query complexity well but still fail as governance?
- Can semantic routing couple similarity matching with resource constraints?
- What makes routing a governance mechanism rather than just a predictor?
- Does model capability still matter once coordination infrastructure is optimized?
- Does small heterogeneous model architecture outperform large homogeneous pools economically?
- How do KNN and prompted routers differ in the accuracy-stability tradeoff?
- What makes mixture-of-experts routing learn token-level specialization effectively?
- Does surface-form query rewriting allow attackers to steer model routing decisions?
- Can hierarchical vector routing reduce context overhead while maintaining tool coverage?
- How do routers decide when to escalate from small to large models?
- Why do KNN routers collapse when queries are paraphrased?
- Why does Branch-Train-Merge fail without learned routing between experts?
- How does semantic clustering help decide which model handles each query?