Line of inquiry
Inquiring lines›How do we ensure safety, alignment…›How can effective AI defenses with…›this line of inquiry
Can single-point security defenses protect multi-agent systems from multi-step attacks?
A broader line of inquiry — a family of 70 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 70
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Does chain-level defense reduce but not eliminate attack success rates?
- Should defense against coordinated intrusion span multiple execution episodes?
- How does prompt hardening work differently in single-agent versus multi-agent systems?
- How do hardened prompts defend against adversarial attacks in multi-agent systems?
- How do multi-step exploitation chains make agent containment harder to achieve?
- Does prompt hardening equally protect single and multi-agent web systems?
- Can defenses at planning boundaries catch attacks that bias upstream instruction signals?
- Can a single security protection work across different system architectures?
- Can input-boundary defenses guard unmonitored channels between agent hops?
- Why do single-boundary defenses underperform in multi-agent systems?
- Can adversarial attacks chain multiple skills to evade security checks?
- Do per-hop inspection gates miss attacks that bias upstream planning signals?
- Should input defenses be validated separately for each channel?
- Why do single-agent and multi-agent systems show different defense effectiveness?
- Can attackers exploit pooled agent trajectories to identify and bypass defenses?
- Can prompt hardening reduce signal propagation in multi-agent systems?
- How do malicious skills evade detection when composed in specific sequences?
- Why do skill scanners fail when evaluating composed behaviors instead of isolated skills?
- Does amplifying a single-actor failure require different security defenses than preventing it?
- How do defenses that inspect planning signals compare to workflow-level validation?
- How should defenders decide whether to publish detection rules and incident analyses?
- Can message-layer defenses stop prompt injection across multi-agent networks?
- Can action-level attack success rates distinguish contained attacks from prevented ones?
- How do workflow-inspecting defenses fail when contamination enters at planning time?
- What trace-level defenses exist beyond per-step review overhead?
- Can skill scanners detect attacks spanning multiple skills in a chain?
- Can fixed pipelines eliminate planning-time attack surfaces in multi-agent systems?
- Why do input-boundary defenses fail in planner-worker pipelines?
- Can defenses check skill chains at execution time instead of scan time?
- How much does prompt hardening actually defend multi-agent systems?
- Does ChainGuard maintain effectiveness when attackers adapt their approach to the defense?
- What makes planning-time attacks structurally invisible to downstream inspection?
- Does fragmented task decomposition hide malicious objectives from detection?
- How does task decomposition fragment the awareness needed to stop an attack?
- Why do defense metrics fail without specifying the attacker's position?
- Which message channels between agents in pipelines lack input validation?
- Do per-hop channel monitors miss coordinated attacks across multiple message transfers?
- Does withholding interaction history defeat attackers in shared stores?
- How do unmonitored channels between pipeline agents enable security gaps?
- Can an attacker copy a rule that distinguishes trusted agents from compromised ones?
- Can ecosystem-level standards reduce trap detection burden?
- What state-tracking requirements exist for defenses that verify multi-party behavioral invariants?
- Can fixed pipelines eliminate planning-time attacks by sacrificing adaptive coordination?
- Do these five vulnerability classes co-occur in predictable attack sequences?
- How do defenders discover which actions belong to the same coordination episode?
- How do chain-level defenses differ from per-skill scanner detection approaches?
- How can a defense validated on one agent silently fail when the system scales?
- How can detection systems identify loops across sequences of delegations?
- What defensive levers shorten the time before probing gets contained?
- Can existing web security defenses protect agents from content manipulation?
- Where do workflow inspection defenses fail against upstream planning attacks?
- Does terminating an intrusion differ from stopping the agent behind it?
- Why do non-overlapping workloads remain invisible to execution-scoped monitoring?
- What attacks does the agent-specific attack surface decompose into?
- How do the six trap categories map onto detection difficulty?
- Does attack success gap shrink when single-agent baseline is already weak?
- Why does attack generation scale faster than defense engineering?
- What makes a coordination episode the right unit for defense response?
- What makes prospective episode discovery harder than using known group membership?
- Why did the endpoint defender not need attribution to act?
- Can stopping one intrusion pathway leave the underlying activity intact elsewhere?
- How do authorization layers differ from input-boundary defenses in blocking attacks?
- Why does scanning skill pairs not fully prevent cross-skill attacks?
- Do layered defenses work better than single privacy techniques?
- What are the eight attack vectors used to probe agents in OpenART?
- What makes the Telephone Loop attack specific to agent delegation?
- What are the four distinct adversary positions in the A-I-R framework?
- Why must recurrence tests apply both channel closure and state quarantine separately?
- What does recovery mean as a defense contract component?
- What does the five-part defense contract actually require of each part?