Line of inquiry
Inquiring lines›How should agents manage and coord…›How do multi-agent reasoning syste…›this line of inquiry
Can model routing outperform monolithic scaling as an efficiency strategy?
A broader line of inquiry — a family of 24 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 24
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can multiple small models outperform a single large model with good routing?
- What makes routing a better investment than training larger models?
- How do routing and test-time compute scaling work together as optimization axes?
- Can embedding-cluster routing outperform a single frontier model?
- Should model routing decisions account for prompt-tier dependencies?
- Why might diverse smaller models with routing beat one giant model?
- Can routing policies remain meaningful over behaviorally homogeneous model pools?
- Can routing enable heterogeneous SLM-first architectures at scale?
- Can compute allocation and model routing be combined for better results?
- Can routing systems prevent expert models from failing outside their specialty?
- Can model routing and compute allocation work together as independent optimizations?
- How does routing decide between models before generation happens?
- What makes query complexity a better routing signal than response quality?
- Why does single-model routing beat ensemble and cascade approaches on latency?
- How do pre-training and distillation enable minimal routing signals to work?
- Can semantic routing couple similarity matching with resource constraints?
- Can a router predict query complexity well but still fail as governance?
- Does small heterogeneous model architecture outperform large homogeneous pools economically?
- Does model capability still matter once coordination infrastructure is optimized?
- What makes mixture-of-experts routing learn token-level specialization effectively?
- Can hierarchical vector routing reduce context overhead while maintaining tool coverage?
- How do routers decide when to escalate from small to large models?
- Why does Branch-Train-Merge fail without learned routing between experts?
- How does semantic clustering help decide which model handles each query?