How do we contain capable agents during evaluation?
Capability tests and attack catalogs exist separately, but little guidance addresses how to keep a powerful agent bounded within its testing environment. This gap matters because evaluation containment is where safety and capability measurement meet.
The abstract states the gap in one sentence: "Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environments used to evaluate it." The introduction adds why the boundary is hard to study: "the evidence needed to study it is scattered: agent-security research, cyber-capability evaluations, containment work, and incident reports each hold a piece of it."
Two features of the claim deserve attention. First, the hedge: "less guidance," not none. The review is not claiming the topic is untouched, only that the pieces sit in separate literatures. Second, the structure of the gap. Capability evaluation asks how strong an agent is. Attack catalogs ask how an agent's components get hit. Neither asks what surrounds a strong agent while it is being tested, so the question falls between them. That makes the review's contribution assembly: it collects what four bodies of work each hold a piece of and organizes it into five vulnerability classes.
The vault reading is that this is one more case of evidence existing but being organized by disciplinary habit. Why do safety failures remain invisible to our evaluation methods? makes the same kind of argument about safety failures generally. It is a different gap from the one in Do cybersecurity benchmarks actually measure exploitation?, where a step of the attack chain is under-measured; here the under-served object is the environment around the measurement.
A second paper places a gap of the same shape at the same object. Do current reward-hacking defenses provide reusable evidence of safety? names a different missing piece, reusable per-run evidence that a run stayed inside its boundary, and there the boundary is the reward path and not the containment of the agent. Both are the authors' own positioning statements with no survey behind them in the excerpts, so together they show two sets of authors locating a gap at the evaluation boundary and do not show that the gap exists.
Caveat. This is the review authors' description of the literature. The excerpt shows no survey, so the vault cannot verify the gap from what it holds.
Inquiring lines that read this note 9
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do evaluation methodologies affect which model capabilities are revealed or hidden?- What makes an evaluation environment itself a security boundary?
- How does evaluation environment design become part of the security boundary?
- Can evaluation environments themselves become security exposures during capability testing?
- How do four separate fields each hold pieces of evaluation safety?
- Is the evaluation environment itself part of the security boundary?
- Can evaluation environments contain security boundaries if they hold shared resources?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do safety failures remain invisible to our evaluation methods?
Current evaluation practices assume failures are obvious, localized, and immediate. But as AI systems deploy into workflows, failures are becoming quiet, distributed, and normalized before detection. What blindspots does this mismatch create?
the general form of "evidence exists but sits where no one looks"
-
Do cybersecurity benchmarks actually measure exploitation?
Frontier models score well on vulnerability finding, patching, and CTF challenges, but does that success tell us whether they can convert vulnerabilities into real attacks? The paper argues exploitation—turning a bug into actual impact—remains under-evaluated.
a neighboring gap in the same field: a missing measurement, where this one is missing containment guidance
-
What failure modes emerge when agents operate without direct oversight?
When autonomous agents are deployed with tool access and memory but without real-time owner oversight, what kinds of failures occur at the agentic layer itself? Understanding these patterns matters for safe deployment.
a component-level failure catalog of the kind the review says leaves containment out
-
Is your evaluation environment actually part of the threat model?
When AI systems can act through tools and credentials during testing, does the evaluation setup itself become a security risk? This explores whether capability measurement and containment are inseparable.
the systems conclusion the assembled evidence is used to support
-
Do current reward-hacking defenses provide reusable evidence of safety?
Existing defenses against reward hacking—task-specific patches, prompt instructions, and post-hoc detectors—may work in practice, but do they leave behind portable, per-run evidence that an evaluation stayed within its intended boundary?
a sibling gap claim about the same object from the reward side: what is missing there is per-run evidence, not containment guidance
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- Agents of Chaos
- Open-World Evaluations for Measuring Frontier AI Capabilities
- Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
- SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems
- Securing Agentic AI: From Per-Action Checks to Trajectory Assurance
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
- Artifacts as Memory Beyond the Agent Boundary
Original note title
existing work separately measures cyber capability and catalogs attacks against agent components — and provides less guidance on containing a capable agent within the environments used to evaluate it