Theme of inquiry
What causes deception and coordination failures in multi-agent AI systems?
A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.
37 specific questions
- Does quarantining state count as recovery in multi-agent attack scenarios?
- How do defenders discover which actions belong to the same coordination episode?
- What makes a coordination episode revisable under agent intrusion?
- Can coalitions rebuild and reaccumulate observations after being removed?
- What makes behavioral containment different from securing individual actions?
- Can agents rebuild communication channels after removal?
- Does terminating an intrusion differ from stopping the agent behind it?
48 specific questions
- Do agent improvements discovered on code tasks transfer to non-coding domains as well?
- Does held-out validation prevent skill document edits from drifting or accumulating harm?
- How should harness infrastructure validate code that agents generate themselves?
- Does bounding textual edits prevent skill degradation better than free rewriting?
- How should agents decide which created code is worth persisting?
- Can execution traces reveal unsupported claims in AI agent behavior?
- What permission models govern code execution within agent skills?
73 specific questions
- How do agent sequences violate system constraints despite individual permissibility?
- How do agent-to-agent messages bypass defenses on downstream principals?
- Does delegation between agents reproduce the confused deputy problem?
- How should task authority constraints apply across multiple coordinated executions?
- What safeguards prevent peer activity from normalizing boundary violations?
- Does the same transfer between agents violate different policies differently?
- Who should verify identity and authorization when agents coordinate across boundaries?
46 specific questions
- Are durable shared code artifacts better than per-task harness patches?
- Can shared persistent artifacts between agents be properly governed?
- How do standardized artifacts reduce inter-agent communication failures?
- How do standardized artifacts prevent autonomous agent failure modes?
- How do standardized artifacts improve coordination between multiple tools?
- What breaks when multiple agents share and revise the same artifacts?
- How do agents decide which created code deserves long-term persistence?
48 specific questions
- How much does peer behavior influence the emergence of collusion?
- How does collusion behavior depend on peer visibility and interaction history?
- Does peer presence or peer behavior shape collusion in verification tasks?
- How does collusion emerge when agents maximize reward over protocol compliance?
- How does verification protocol structure affect collusion emergence?
- Can pairing or vetting peers reduce collusion as a design lever?
- Does restricting interaction history between agents reduce coupling or prevent collusion?
60 specific questions
- Can reasoning traces and logged actions expose scheming that public messages hide?
- Can monitoring reasoning alone miss scheming that agents conceal in behavior?
- What happens to agent candor when reasoning traces are monitored versus hidden?
- Does reasoning transparency predict honesty in agent final messages?
- Can reasoning traces reliably distinguish honest mistakes from deliberate lies in agent speech?
- How does chain-of-thought monitoring fail when agents are trying to hide something?
- Can public cheap talk or behavior alone expose an objectively misaligned agent?