SYNTHESIS NOTE
Topics›Agent Harness›this note

Does measuring exploit capability help or harm defense?

Exploitation benchmarks can support defenders and attackers equally. How should we evaluate capabilities with unavoidable dual-use potential, and what safeguards make evaluation itself defensible?

Synthesis note · 2026-09-23 · sourced from Agent Harness

The ExploitGym abstract makes a short but consequential statement: exploitation "is inherently dual-use, supporting defensive workflows while lowering the barrier for offense." It is presented as a property of the capability itself, not of how someone chooses to use it. The excerpt gives no example of the defensive side. As an illustration of my own, an agent that can extend a vulnerability into a working exploit could help a defender show that a bug is exploitable and worth fixing first, and it could equally help someone with less skill do the same for the wrong reason.

Two consequences follow. First, a single measurement of the capability cannot be labeled as either good news or bad news. A higher score means better defensive tooling and greater misuse potential at once, so the interpretation depends on who has access and under what controls. Second, the argument for evaluating it anyway is the paper's own: the authors describe exploitation as important, diagnostic and "under-evaluated," and treat rigorous evaluation as urgent. The evaluation is justified by the risk the capability creates, even though the evaluation is itself close to the capability.

This is the same shape as other dual-use findings in the vault. Can AI reduce conspiracy beliefs by tailoring counterevidence personally? shows a persuasion capability working in the prosocial direction, while the persuasion notes on the risk side show the same mechanism working against people. Exploitation adds a case where the dual-use ambiguity is stated up front by the benchmark's own authors. And it bears on the Can we measure how much risk open models actually add? question: "lowering the barrier for offense" is a marginal-risk claim, and an exploitation benchmark is one way to start measuring the barrier instead of asserting it.

How close the evaluation sits to the capability is not only a matter of interpretation. The cyber-capable-agents review argues that once an agent has execution access, the environment it is evaluated in is inside the security boundary (Is your evaluation environment actually part of the threat model?), so a run is an exposure event as well as a measurement. That pulls against the urgency above, and the vault has filed it as a tension; it probably dissolves if evaluation scales only as fast as containment does, a vault hypothesis that neither paper states.

The excerpt does not say how ExploitGym handles its own release, access or containment of its runs, so nothing here should be read as a claim about that.

Inquiring lines that read this note 14

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How prevalent is reward hacking in frontier models? Do AI capability benchmarks accurately measure reasoning ability or just surface patterns? Do planted honeypot tests reliably measure reward hacking? Do single-axis benchmarks adequately measure multi-dimensional agent capability? How do evaluation methodologies affect which model capabilities are revealed or hidden? Do current AI defenses adequately protect against semantic manipulation attacks? Do frontier models develop hidden self-protective behaviors? How does outcome-only reporting obscure which system components blocked attacks?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 119 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

exploitation is inherently dual-use — the same capability that supports defensive workflows lowers the barrier for offense