Theme of inquiry
What causes coordination failures and safety problems in multi-agent systems?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
66 specific questions
- Do single-agent systems outperform multi-agent coordination as model capabilities grow?
- When does multi-agent scaling actually outperform static ensembles?
- Do specialized agents outperform single agents with better orchestration?
- How do multi-agent systems improve on single frontier models?
- Which research tasks are better suited for multi-agent versus single-agent approaches?
- Which failure mode most limits current multi-agent performance?
- Can multi-agent teams solve problems better than single models thinking longer?
19 specific questions
- How do tool invocations drive agentic cost beyond token consumption?
- Can latent communication reduce the token cost of multi-agent systems?
- When is 15x token overhead actually worth the compute cost?
- Do multi-agent systems justify their token costs with genuine quality gains?
- Does effective feedback compute matter more than raw token expenditure for agent scaling?
- Does upgrading model capability improve token efficiency in agentic systems?
- Why do multi-agent systems use 15 times more tokens than chat interactions?
66 specific questions
- Can multi-agent debate prevent the confident convergence on wrong answers?
- Why do multi-agent systems converge on wrong answers without debate safeguards?
- Can multi-agent debate prevent reasoning models from amplifying errors?
- What mechanisms drive silent agreement in multi-agent reasoning systems?
- How often do AI agents reach false agreement in group reasoning tasks?
- Why does premature consensus form in multi-agent reasoning systems?
- Why does premature consensus form in multi-agent reasoning without genuine deliberation?
41 specific questions
- How do multi-agent LLM systems fail at coordination and role consistency?
- Do multi-agent language model teams fail the same way individual reasoning does?
- Do multi-agent LLM systems fail in measurably different ways than single agents?
- Why do multi-agent LLM systems converge prematurely without genuine deliberation or probing?
- Does multi-agent interaction amplify existing failures or create new ones?
- Why does silent agreement occur so often in multi-agent LLM systems?
- How do LLM-based agents develop shared abstractions through interaction?
63 specific questions
- How do standardized artifacts prevent autonomous agent failure modes?
- Do agent-created languages improve or degrade performance on their original tasks?
- How do standardized artifacts reduce inter-agent communication failures?
- What would unified agent-to-agent and agent-to-tool protocols actually look like?
- What role does standardization play in multi-agent system ecosystems?
- How do standardized artifacts improve coordination between multiple tools?
- What makes protocols better than free-form prompting for tool coordination?