SYNTHESIS NOTE
Topics›Alignment›this note

How do we contain capable agents during evaluation?

Capability tests and attack catalogs exist separately, but little guidance addresses how to keep a powerful agent bounded within its testing environment. This gap matters because evaluation containment is where safety and capability measurement meet.

Synthesis note · 2026-09-23 · sourced from Alignment

The abstract states the gap in one sentence: "Existing work separately measures cyber capability and catalogs attacks against agent components, but provides less guidance on containing a capable agent within the environments used to evaluate it." The introduction adds why the boundary is hard to study: "the evidence needed to study it is scattered: agent-security research, cyber-capability evaluations, containment work, and incident reports each hold a piece of it."

Two features of the claim deserve attention. First, the hedge: "less guidance," not none. The review is not claiming the topic is untouched, only that the pieces sit in separate literatures. Second, the structure of the gap. Capability evaluation asks how strong an agent is. Attack catalogs ask how an agent's components get hit. Neither asks what surrounds a strong agent while it is being tested, so the question falls between them. That makes the review's contribution assembly: it collects what four bodies of work each hold a piece of and organizes it into five vulnerability classes.

The vault reading is that this is one more case of evidence existing but being organized by disciplinary habit. Why do safety failures remain invisible to our evaluation methods? makes the same kind of argument about safety failures generally. It is a different gap from the one in Do cybersecurity benchmarks actually measure exploitation?, where a step of the attack chain is under-measured; here the under-served object is the environment around the measurement.

A second paper places a gap of the same shape at the same object. Do current reward-hacking defenses provide reusable evidence of safety? names a different missing piece, reusable per-run evidence that a run stayed inside its boundary, and there the boundary is the reward path and not the containment of the agent. Both are the authors' own positioning statements with no survey behind them in the excerpts, so together they show two sets of authors locating a gap at the evaluation boundary and do not show that the gap exists.

Caveat. This is the review authors' description of the literature. The excerpt shows no survey, so the vault cannot verify the gap from what it holds.

Inquiring lines that read this note 9

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do evaluation methodologies affect which model capabilities are revealed or hidden? Can defenses detect attacks composed across multiple skills? How do coordinated agent sequences violate constraints that individual actions respect? How can defenders detect coordinated attacks across episodes?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 117 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

existing work separately measures cyber capability and catalogs attacks against agent components — and provides less guidance on containing a capable agent within the environments used to evaluate it