Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›What causes coordination failures…›this line of inquiry
When do multi-agent systems outperform single frontier models?
A broader line of inquiry — a family of 66 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 66
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do single-agent systems outperform multi-agent coordination as model capabilities grow?
- When does multi-agent scaling actually outperform static ensembles?
- Do specialized agents outperform single agents with better orchestration?
- How do multi-agent systems improve on single frontier models?
- Which research tasks are better suited for multi-agent versus single-agent approaches?
- Which failure mode most limits current multi-agent performance?
- Can multi-agent teams solve problems better than single models thinking longer?
- At what capability threshold does multi-agent coordination stop helping?
- How do context engineering limits relate to multi-agent coordination problems?
- At what task difficulty does multi-agent decomposition become worth the coordination cost?
- Which layer of agent systems creates the largest capability gains in practice?
- When does multi-agent routing quality actually exceed single-agent or static ensemble performance?
- Does parallel task structure determine optimal multi-agent architecture?
- Can cognitive diversity compensate for lack of expertise in agent teams?
- How does role allocation in multi-agent systems depend on model differentiation?
- Does horizontal coordination improve with stronger individual agents?
- What four decisions matter most in multi-agent system routing?
- Does internal task decomposition eliminate overhead from multi-agent coordination?
- Why do comparable metrics matter across different multi-agent system designs?
- How do multi-agent routers balance flexibility against interpretability in design?
- When does multi-agent voting help versus hurt performance on tasks?
- Can cognitive diversity overcome expertise gaps in agent teams?
- How does distributed coordination fail as agent networks scale?
- What quantitative costs and failure modes emerge when coordinating multiple agents?
- Do multi-agent LLM systems scale better than centralized hierarchies?
- How does collaboration topology choice affect error amplification in multi-agent systems?
- How does role specialization preserve reasoning diversity in multi-agent teams?
- How do static team decomposition and dynamic agent selection compare in efficiency?
- Can designated leadership structures reduce premature convergence in multi-agent reasoning?
- How do capability vectors enable discovery in multi-agent systems?
- What interaction topologies and agent counts best exploit complementary expertise?
- Can small language models handle diverse tasks in heterogeneous multi-agent systems?
- Why does diversity without expertise produce worse results than a single capable agent?
- Can one routing framework cover topology and role choices in multi-agent systems?
- Is the coupled human-agent environment the right unit for evaluation?
- What distinguishes collective evolution from vertical self-improvement in agent systems?
- Can a single manager policy work across vastly different agent architectures?
- How much does confidence-guided cascading between SAS and MAS improve accuracy?
- How does multi-agent reasoning scale compared to single-model approaches?
- How do we measure coordination when multiple agents act together?
- What equilibrium-selection problem does human data solve in multi-agent learning?
- Does cognitive diversity in teams only pay off when agents actively explore it?
- How do cognitive stimulation and process losses interact in group AI systems?
- Why does capability discovery become the bottleneck in large agent systems?
- How do agent behaviors aggregate into prices and allocations?
- How will the agent economy reshape compute infrastructure design?
- What makes capability vectors a better coordination substrate than topic-based routing?
- Why does literature review benefit most from multi-agent orchestration approaches?
- How does network structure affect whether agent communities improve or amplify collective reasoning?
- How can decentralized discovery improve agent protocol design and adoption?
- What structural features drive instrumental convergence across different agent goals?
- Can we decompose agent efficiency into measurable independent components?
- How does coordination governance shift the hard problem from capability itself?
- Which ecosystem conditions matter most for agent deployment success?
- How can controlled experiments isolate multi-agent interaction effects from architecture?
- What five ecosystem conditions must coordination governance and evidence actually satisfy?
- Can multi-agent reasoning systems scale beyond current architectures?
- What ecosystem conditions must exist for agents to function as economic participants?
- Why does decentralization work better than central planning for open-ended research?
- What ecosystem conditions make agent attention markets viable?
- How do sharded HNSW indices preserve capability distinctions at scale?
- What capability threshold do agents need to self-organize effectively?
- What causes prolonged concentration on single approaches in decentralized research teams?
- What distinguishes a component's link to the collective from coupling among defecting components?
- When should multi-agent systems escalate rather than aggregate toward a single decision?
- How do controllable simulators compare to population-level agent simulation approaches?