Line of inquiry
Inquiring lines›How do agents behave and coordinat…›How do agents in production pipeli…›this line of inquiry
Do multi-agent systems create greater security risks than single-agent ones?
A broader line of inquiry — a family of 47 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 47
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- When do multi-agent architectures create more attack surface than single-agent systems?
- How does task division in multi-agent design affect security outcomes?
- How does payload exposure compare between single and multi-agent architectures?
- How does a single compromised agent degrade performance across entire multi-agent pipelines?
- How does prompt hardening work differently in single-agent versus multi-agent systems?
- Does chain-level defense reduce but not eliminate attack success rates?
- What makes agent-to-agent messages in multi-agent systems vulnerable to exploitation?
- Does prompt hardening equally protect single and multi-agent web systems?
- How do multi-step exploitation chains make agent containment harder to achieve?
- Can specialized roles let malicious objectives hide across multiple agents?
- Does the architectural penalty of MAS hold across different models and scenarios?
- Can single-agent defenses prevent cascading failures in multi-agent systems?
- Does amplifying a single-actor failure require different security defenses than preventing it?
- Can safe individual AI agents fail when deployed together?
- Why do single-agent and multi-agent systems show different defense effectiveness?
- Do server-side filters hide the true success rate of multi-agent attacks?
- Why do single-boundary defenses underperform in multi-agent systems?
- What attacks are unique to multi-agent systems compared to single agents?
- Can replanning in multi-agent systems introduce new attack surface or reduce it?
- Can shared memory poisoning compromise multi-agent delegation chains?
- Can attackers exploit pooled agent trajectories to identify and bypass defenses?
- How does task decomposition hide harmful objectives across multiple agents?
- Can adversarial attacks chain multiple skills to evade security checks?
- Can task decomposition allow harmful objectives to hide in locally plausible subtasks?
- How do malicious skills evade detection when composed in specific sequences?
- What baseline would prove multi-agent systems are actually less safe?
- Does quarantining state count as recovery in multi-agent attack scenarios?
- Does fragmented task decomposition hide malicious objectives from detection?
- Do server-side filters hide the true strength of multi-agent attacks?
- Which interaction interfaces do multi-agent systems expose to adversaries?
- Which agent properties like state retention enable supply-chain and credential vulnerabilities?
- How do single-agent safety evaluations underestimate risks in deployed multi-agent systems?
- How much does prompt hardening actually defend multi-agent systems?
- Does withholding interaction history defeat attackers in shared stores?
- What containment risks emerge as agents obtain successive exploit primitives?
- Can ecosystem-level standards reduce trap detection burden?
- How does task decomposition fragment the awareness needed to stop an attack?
- Do these five vulnerability classes co-occur in predictable attack sequences?
- Can an attacker copy a rule that distinguishes trusted agents from compromised ones?
- Why does vulnerability to extortion actually promote cooperation between agents?
- Does attack success gap shrink when single-agent baseline is already weak?
- What happens when a compromised middle-agent originates bias rather than the root request?
- What attacks does the agent-specific attack surface decompose into?
- What mitigation strategies prevent misaligned agents from harming team outcomes?
- How does insider threat differ from external attack in multi-agent systems?
- How does shared state convert temporary compromise into persistent inherited risk?
- What are the four distinct adversary positions in the A-I-R framework?