Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›What security vulnerabilities enab…›this line of inquiry
Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid?
A broader line of inquiry — a family of 46 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 46
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can safe individual AI agents fail when deployed together?
- What makes agent-to-agent messages in multi-agent systems vulnerable to exploitation?
- How does task division in multi-agent design affect security outcomes?
- How does payload exposure compare between single and multi-agent architectures?
- When do multi-agent architectures create more attack surface than single-agent systems?
- How does a single compromised agent degrade performance across entire multi-agent pipelines?
- What baseline would prove multi-agent systems are actually less safe?
- Can single-agent defenses prevent cascading failures in multi-agent systems?
- Can shared memory poisoning compromise multi-agent delegation chains?
- How do multi-agent systems fail when agents cannot verify each other's claims?
- Can multi-agent architecture isolation reveal which design choices matter most for safety?
- How do single-agent safety evaluations underestimate risks in deployed multi-agent systems?
- Does the architectural penalty of MAS hold across different models and scenarios?
- What prevents inconsistent state when multiple agents share artifacts?
- What breaks when multiple agents share and revise the same artifacts?
- Can a single safe model guarantee safety in multi-agent composition?
- Why do multi-agent failures arise through interactions local checks miss?
- Can specialized roles let malicious objectives hide across multiple agents?
- What attacks are unique to multi-agent systems compared to single agents?
- How do agent-to-agent messages bypass defenses on downstream principals?
- How does task decomposition hide harmful objectives across multiple agents?
- What are the four mechanisms that carry failures across agent boundaries?
- Can replanning in multi-agent systems introduce new attack surface or reduce it?
- Can task decomposition allow harmful objectives to hide in locally plausible subtasks?
- Where should the trust boundary sit in multi-agent planning systems?
- Which interaction interfaces do multi-agent systems expose to adversaries?
- How do shared artifact stores become security risks in multi-agent systems?
- Does quarantining state count as recovery in multi-agent attack scenarios?
- Who can actually observe and challenge errors in multi-agent AI workflows?
- Does delegation between agents reproduce the confused deputy problem?
- What prevents multiple agents from corrupting shared state in live artifacts?
- Which agent properties like state retention enable supply-chain and credential vulnerabilities?
- Why does monitoring performed by agents on agents create safety risks?
- Can mixed-authorship traces from multi-agent pipelines be monitored reliably?
- What vulnerabilities emerge at each hop between agents in a pipeline?
- Where should the trust boundary sit in multi-agent planner systems?
- What makes unmonitored channels between agents safety-critical?
- What containment risks emerge as agents obtain successive exploit primitives?
- Why does agent-to-agent interaction expose identity verification vulnerabilities?
- What baseline comparison shows whether interaction actually caused multi-agent failures?
- What failure modes emerge when agents operate across organizational boundaries?
- How do agents inherit exploit knowledge through shared history?
- What governance structures prevent harmful coordination as AI agents multiply?
- How does insider threat differ from external attack in multi-agent systems?
- Why are unmonitored channels between agents a safety risk?
- Can protocol bridges introduce new failure modes or security vulnerabilities?