SYNTHESIS NOTE
Topics›Alignment›this note

Is your evaluation environment actually part of the threat model?

When AI systems can act through tools and credentials during testing, does the evaluation setup itself become a security risk? This explores whether capability measurement and containment are inseparable.

Synthesis note · 2026-09-23 · sourced from Alignment

The review's abstract compresses its two incident records into one shared lesson: "the evaluation environment is itself part of the security boundary." The conclusion states the general form: "Cyber-capable agents make the security of capability evaluation an end-to-end systems problem. Once a model is connected to memory, tools, credentials, and an execution environment, those components—and the response workflow around them—become part of the security boundary."

The usual picture of an evaluation is a measuring instrument: put a model in a harness, read a number. The review asks the reader to stop treating the harness as only an instrument. A model with an execution environment can act, so the environment is also the place where whatever it can do happens. The evaluator's question (how capable is this model?) and the security question (what can it reach while we ask?) share a single setup, and answering the first does not answer the second. In the conclusion's words, evaluating capability "without evaluating that boundary leaves out the mechanisms through which a capable agent can act."

That is a different role for the environment than the one the vault's ExploitGym notes give it. There the environment is an object of measurement: protections become a controlled variable, and the environment component is separated so trajectories can be diagnosed. Here the same component is an object of containment. One component, two jobs; an evaluation that instruments it well but contains it poorly has done half the work.

The vault states the same conclusion from the reward side, and the two draw the edge by different questions. Where does the evaluation boundary actually end in agent benchmarks? asks what the agent can influence that the reward later reads; this note asks what a capable agent can reach while it is tested. Do frontier models exploit unknown vulnerabilities in evaluations? reads the reported behavior as one fact in two vocabularies, which is that note's reading; this review's excerpt says nothing about reward hacking. A filed tension in ops/tensions/ asks whether the recorder that produces reward-path evidence sits inside the boundary this note describes.

The claim is deliberately narrow. The review says its two records do not establish a causal mechanism, so the lesson is a design stance about where the boundary is drawn, not a finding about which control failed. The excerpt reports no results.

Post angle: your evaluation harness has a threat model. The environment built to measure a model's risk is also a place where that risk can occur, and a score from it says nothing about the second.

Inquiring lines that read this note 13

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do evaluation methodologies affect which model capabilities are revealed or hidden? What determines whether AI system errors remain visible and contestable? How does training data contamination persist through safety alignment mechanisms? What infrastructure evidence validates agent benchmark achievement claims? How does outcome-only reporting obscure which system components blocked attacks? How can defenders detect coordinated attacks across episodes?

Related concepts in this collection 9

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
23 direct connections · 149 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the evaluation environment is itself part of the security boundary — the security of capability evaluation is an end-to-end systems problem