SYNTHESIS NOTE
Topics›Evaluations›this note

How representative is the BenchShield Trajectories labeled sample?

The corpus contains 456 human-labeled trajectories from over 31,000 public runs—about 1.5 percent. Whether this subset can estimate actual reward hacking rates depends entirely on how those 456 were selected, a choice the paper does not disclose.

Synthesis note · 2026-09-24 · sourced from Evaluations

The abstract's last sentence: "We construct BenchShield Trajectories, a human-labeled corpus of 456 adjudicated trajectories from more than 31,000 public agent runs across three benchmarks." The excerpt's abstract stops there. It reports no detection performance, no hack rate and no label counts, so whatever the paper shows on the corpus is outside what is available here.

By vault arithmetic, 456 of more than 31,000 is at most about 1.5 percent of the runs. Whether a rate can be read from such a subset depends on how it was drawn. A random sample supports an estimate of how often public agent runs reward hack, with an error that depends on the number labeled. A sample drawn toward suspicious runs supports no rate at all but is a useful test bed with hacks in it. "Adjudicated" suggests labels were resolved after disagreement or a second review, which says something about label quality and nothing about selection.

The interest for the vault is the source of the runs. They are "public agent runs," made for other purposes, not runs built around bait. Most of the vault's rates are under planting: How often do frontier agents exploit planted reward hacking shortcuts? is a rate under an optional shortcut, and Does BaitBench measure hacking propensity or bait visibility? asks what such a rate measures. The exception is How often do models hack unmodified coding benchmarks?, a rate for one model on two standard benchmarks with no planted shortcut mentioned, though the excerpt does not say whose runs those were or how a hack was labeled. A rate from public runs across three benchmarks, with labels a human adjudicated, would be a second one without the planting caveat, if the corpus supports one.

The corpus is also a label source. "Human-labeled" and "adjudicated" are one answer to the question How were reward hacks labeled in this benchmark study? puts to another paper's numbers, and the same test applies here: a human label is a reference only as far as its procedure is stated. The abstract gives no adjudication rule, no label counts and no agreement figure.

What would answer it: the sampling rule for the 456, the label distribution, and how BenchShield's output agrees with the human labels. The excerpt gives none.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does outcome-only reporting obscure which system components blocked attacks?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 79 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

does the BenchShield Trajectories corpus estimate how often agents reward hack — the excerpt gives 456 adjudicated trajectories from more than 31000 public runs and does not say how the 456 were chosen