Theme of inquiry
How do agents in production pipelines exploit and bypass operational constraints?
A question within its area, explored through 4 lines of inquiry below — each a family of specific questions the research asks.
29 specific questions
- Can we build reusable evidence that a run stayed within bounds?
- Can infrastructure evidence ground benchmark claims better than terminal scores alone?
- Why does infrastructure-side evidence matter more than agent-reported traces?
- Which reset, logging, and feedback channels do agents actually exploit in benchmarks?
- How can operators ground benchmark completion claims in infrastructure data?
- Can task success alone reveal whether memory routing is working?
- What evidence should benchmark operators attach to completion claims?
47 specific questions
- When do multi-agent architectures create more attack surface than single-agent systems?
- How does task division in multi-agent design affect security outcomes?
- How does payload exposure compare between single and multi-agent architectures?
- How does a single compromised agent degrade performance across entire multi-agent pipelines?
- How does prompt hardening work differently in single-agent versus multi-agent systems?
- Does chain-level defense reduce but not eliminate attack success rates?
- What makes agent-to-agent messages in multi-agent systems vulnerable to exploitation?
62 specific questions
- Can skill repositories evolve toward execution-oriented refinement over time?
- Are durable shared code artifacts better than per-task harness patches?
- How do agents decide which created code deserves long-term persistence?
- Where does agent reliability come from if not better tools?
- How do agents decide which created code should persist versus disappear?
- Do learned workflows transfer between different agents with minimal accuracy loss?
- How do externalizing cognitive work and coordination infrastructure relate to agent reliability?
46 specific questions
- Do agents systematically misreport their own capabilities and tool access?
- Which interaction artifacts matter most for reliable agent evaluation?
- What makes recorded transitions more trustworthy than agent reasoning trajectories?
- How should harness infrastructure validate code that agents generate themselves?
- Do agents interpret peer edits as legitimate prior changes versus tampering?
- How should we label ground truth when a protected state change alone is ambiguous?
- Why does forcing agents to trace function paths prevent unsupported claims?