Why do AI security tests grade the steps around an attack but rarely measure the one that turns a bug into a working exploit?
Why is exploitation the most under-measured stage in cybersecurity benchmarking?
This explores why cybersecurity benchmarks for AI rarely test exploitation (the step where a known vulnerability is turned into a working attack) even though they test the stages before and after it, and what makes that step hard to measure.
This explores why AI cybersecurity benchmarks skip the step where a bug becomes a working attack, even though they measure the stages around it. The corpus says this directly: frontier models are already tested and do well at reproducing vulnerabilities, writing patches and solving capture-the-flag puzzles. Exploitation, the step that makes a vulnerability dangerous, is still largely unmeasured Do cybersecurity benchmarks actually measure exploitation?. So a strong score on today's security benchmarks doesn't tell you whether a model can carry an attack through.
The corpus suggests a few reasons for the gap. The first is not technical. Exploitation is dual-use: the skill that helps a defender judge how serious a bug is also makes attacks easier for an attacker. A score can't tell you which of those you're measuring. That depends on who has access to the model and under what controls Does measuring exploit capability help or harm defense?. Publishing a good exploitation benchmark therefore also publishes a measure of offensive capability, which may help explain why builders have avoided it.
The second reason is the scoring itself. ExploitGym counts success as arbitrary code execution, meaning the attacker can run any code it wants on the target. That is easy to check and clearly severe. But real exploits are built in stages, such as gaining the ability to read or write arbitrary memory or escaping a sandbox. A pass/fail endpoint scores an agent that got most of the way there the same as one that failed at the start Does arbitrary code execution alone capture exploit progress?. A benchmark that only records the final success can hide steady progress until the day the score jumps.
The third reason is less obvious: there are few published answer keys. Complete working exploits rarely appear in public, so benchmark builders lack reference solutions. That same scarcity also helps the benchmark, because a model can't have memorized answers that were never published. It has to build the exploit itself. That protection fades once solutions are published after the benchmark appears Can scarcity of solutions protect benchmarks from data contamination?. The thing that makes exploitation hard to benchmark is also what makes a benchmark of it trustworthy, at least for a while.
The corpus also uses "exploit" in a second sense: agents gaming the evaluation itself. Here, researchers have separated tasks that merely leave an attack path open from runs where the agent actually used it. They do this by recording key actions at the infrastructure level as the agent runs Can runtime instrumentation distinguish hacking exposure from actual exploitation?, and by checking each run against a model of what the task should involve Can a finite lifecycle model detect reward hacking across benchmarks?. These papers aren't about cyberattacks. Still, recording what an agent actually did, rather than only whether it reached the endpoint, is the kind of method that could also score partial exploit progress. The corpus doesn't yet contain work that applies it to cybersecurity exploitation, so that link is an open question, not a finding.
Sources 6 notes
ExploitGym shows that while frontier models excel at vulnerability reproduction, patch generation, and CTF tasks, exploitation—the step where a vulnerability becomes a real attack—remains largely unmeasured in the benchmark literature.
ExploitGym demonstrates that exploit generation supports both defensive vulnerability assessment and lowering barriers to offensive attacks simultaneously. No single measurement distinguishes between these outcomes without knowing who has access and under what controls.
ExploitGym's success criterion—arbitrary code execution—is verifiable and clear but ignores meaningful outcomes like arbitrary read/write primitives and sandbox escape. This endpoint-only metric treats agents that reach intermediate steps identically to those that fail immediately.
ExploitGym's missing ground-truth exploits reduce data contamination risk because complete working exploits are not widely published. Models must construct solutions rather than recall them, though this protection may erode as solutions are published post-benchmark.
Infrastructure-side recording of authority-bearing transitions distinguishes tasks that merely expose a hacking vector from runs that actually exercise one. This separation prevents every score from an exposed task being automatically suspect.
Show all 6 sources
BenchShield grounds exploit detection in a finite lifecycle model of reward-relevant events, enabling both static pre-run analysis and runtime instrumentation to operate over the same formal object. This replaces pattern-matching detection with deviation from an intended lifecycle structure.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?
- Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks
- OpenAI and Hugging Face partner to address security incident during model evaluation
- ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
- Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks
- Recent Frontier Models Are Reward Hacking
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
- Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations