Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do adversarial attacks exploit…›this line of inquiry
How do multi-agent architectures affect AI system security and defense effectiveness?
A broader line of inquiry — a family of 83 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 83
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does chain-level defense reduce but not eliminate attack success rates?
- How do hardened prompts defend against adversarial attacks in multi-agent systems?
- Does prompt hardening equally protect single and multi-agent web systems?
- When do multi-agent architectures create more attack surface than single-agent systems?
- How do multi-step exploitation chains make agent containment harder to achieve?
- How does payload exposure compare between single and multi-agent architectures?
- How does task division in multi-agent design affect security outcomes?
- Can defenses at planning boundaries catch attacks that bias upstream instruction signals?
- How does prompt hardening work differently in single-agent versus multi-agent systems?
- Can prompt hardening reduce signal propagation in multi-agent systems?
- Do per-hop inspection gates miss attacks that bias upstream planning signals?
- Can a single security protection work across different system architectures?
- Do server-side filters hide the true success rate of multi-agent attacks?
- Does amplifying a single-actor failure require different security defenses than preventing it?
- Can single-agent defenses prevent cascading failures in multi-agent systems?
- Can input-boundary defenses guard unmonitored channels between agent hops?
- Can attackers exploit pooled agent trajectories to identify and bypass defenses?
- How does a single compromised agent degrade performance across entire multi-agent pipelines?
- What makes agent-to-agent messages in multi-agent systems vulnerable to exploitation?
- How does workflow position amplify malicious signals in multi-agent systems?
- Why do single-boundary defenses underperform in multi-agent systems?
- Should defense against coordinated intrusion span multiple execution episodes?
- Can adversarial attacks chain multiple skills to evade security checks?
- Does the architectural penalty of MAS hold across different models and scenarios?
- Why do single-agent and multi-agent systems show different defense effectiveness?
- Why do outcome-level metrics fail to reveal contained attacks in multi-agent pipelines?
- Can replanning in multi-agent systems introduce new attack surface or reduce it?
- Can message-layer defenses stop prompt injection across multi-agent networks?
- What role do server-side filters play in masking multi-agent pipeline vulnerabilities?
- Should input defenses be validated separately for each channel?
- Can specialized roles let malicious objectives hide across multiple agents?
- How do malicious skills evade detection when composed in specific sequences?
- Can fixed pipelines eliminate planning-time attack surfaces in multi-agent systems?
- How should defenders decide whether to publish detection rules and incident analyses?
- Do server-side filters hide the true strength of multi-agent attacks?
- What attacks are unique to multi-agent systems compared to single agents?
- How do defenses that inspect planning signals compare to workflow-level validation?
- Why does workflow position amplify malicious signals in multi-agent relay chains?
- Can action-level attack success rates distinguish contained attacks from prevented ones?
- Do prompt injection attacks propagate behavioral bias across multi-agent networks?
- Does excluding reasoning access leave scheming detectors vulnerable to plan injection attacks?
- Why do skill scanners fail when evaluating composed behaviors instead of isolated skills?
- Which agent properties like state retention enable supply-chain and credential vulnerabilities?
- Which interaction interfaces do multi-agent systems expose to adversaries?
- How do workflow-inspecting defenses fail when contamination enters at planning time?
- How does prompt injection differ from subliminal message propagation in multi-agent networks?
- Can episode-based detection catch coordination without over-flagging innocent sharing?
- Can ecosystem-level standards reduce trap detection burden?
- What makes planning-time attacks structurally invisible to downstream inspection?
- Can fixed pipelines eliminate planning-time attacks by sacrificing adaptive coordination?
- How much does prompt hardening actually defend multi-agent systems?
- Does ChainGuard maintain effectiveness when attackers adapt their approach to the defense?
- Does withholding interaction history defeat attackers in shared stores?
- How do unmonitored channels between pipeline agents enable security gaps?
- Which message channels between agents in pipelines lack input validation?
- Do per-hop channel monitors miss coordinated attacks across multiple message transfers?
- Which backend filters silently affect the reported attack success numbers?
- Do these five vulnerability classes co-occur in predictable attack sequences?
- Why do input-boundary defenses fail in planner-worker pipelines?
- Can skill scanners detect attacks spanning multiple skills in a chain?
- What trace-level defenses exist beyond per-step review overhead?
- Why does workflow position amplify malicious signals downstream?
- Why do defense metrics fail without specifying the attacker's position?
- Can an attacker copy a rule that distinguishes trusted agents from compromised ones?
- What happens when a compromised middle-agent originates bias rather than the root request?
- Can defenses check skill chains at execution time instead of scan time?
- How can a defense validated on one agent silently fail when the system scales?
- What state-tracking requirements exist for defenses that verify multi-party behavioral invariants?
- Can existing web security defenses protect agents from content manipulation?
- How can detection systems identify loops across sequences of delegations?
- How does workflow position amplify or suppress malicious signals?
- Where do workflow inspection defenses fail against upstream planning attacks?
- Does attack success gap shrink when single-agent baseline is already weak?
- How do the six trap categories map onto detection difficulty?
- What attacks does the agent-specific attack surface decompose into?
- Is malicious propagation fundamentally a semantic information flow problem?
- Why does attack generation scale faster than defense engineering?
- Why do non-overlapping workloads remain invisible to execution-scoped monitoring?
- How does insider threat differ from external attack in multi-agent systems?
- Why did the endpoint defender not need attribution to act?
- What are the eight attack vectors used to probe agents in OpenART?
- What are the four distinct adversary positions in the A-I-R framework?
- What injection vectors threaten Kubernetes data-processing pipelines in production?