INQUIRING LINE

A headline number of 1,213 AI incidents sounds solid — but what exactly was counted, and where did they come from?

What population of incidents does the 1,213 count represent?

This explores what the 1,213 figure is a count of: which incidents were collected, from where, and how they were chosen, so a reader knows what the number can and can't speak for.


This explores what the 1,213 figure is a count of: which incidents were collected, from where, and how they were chosen. The corpus can only partly answer that. It says the 1,213 are coded incidents, meaning each one was read and labeled, and that roughly 80% of them contain no record of any stop mechanism, whether technical, operational, legal or third-party How often do incident records document system stops?. It doesn't say what the sampling frame was. There's nothing here about the source database, the time window, or whether the incidents were all reported cases or a selected subset. I won't guess at any of those.

The wording does tell you one thing. The finding is about what the records document, not necessarily what happened. "Four in five record no documented stop" is a claim about incident write-ups, and a write-up that never mentions a stop doesn't prove none existed. So the 1,213 is best read as a population of incident records, and the 80% as a gap in documentation and interruptibility as the paper interprets it.

A nearby note shows why the missing frame matters. A different paper reports 456 adjudicated trajectories out of 31,000+ agent runs but gives no sampling rule, label distribution or adjudication procedure. The note concludes that without those details the sample can't support a defensible estimate of how often agents reward-hack How representative is the BenchShield Trajectories labeled sample?. That is a different corpus, so it says nothing about how the 1,213 were chosen. It does show the standard to apply: a percentage is only as meaningful as the rule that decided which cases got counted. A similar gap turns up elsewhere in the library, where many attack and defense figures don't say whether they were measured behind a filter, so readers can't tell what the number reflects Which attack and defense numbers came from filtered backends?.

Two other notes set expectations for how far incident evidence can be pushed. Analysis of preliminary incident records supports lessons about boundaries but explicitly does not establish recurrence rates or causal mechanisms What can two incident records actually teach us about AI evaluation security?. A comparative protocol keeps what one record claims separate from what several records support together How do you separate reliable claims from fragile early incident evidence?. Applied to the 1,213, that suggests treating the 80% as a strong signal about how incidents get documented, and not as an estimate of how often real systems lack a stop.

If you need the exact population, such as the source, date range or inclusion criteria, the corpus notes here don't contain it. You would need the original paper's methods section.


Sources 5 notes

How often do incident records document system stops?

Analysis of 1,213 coded incidents showed that approximately 80% contain no record of any stop mechanism—technical, operational, legal, or third-party. The paper interprets this frequency as evidence of a significant gap in system interruptibility.

How representative is the BenchShield Trajectories labeled sample?

The paper reports 456 adjudicated trajectories from 31,000+ runs but provides no sampling rule, label distribution, or adjudication procedure. Without these details, the corpus cannot support a defensible estimate of reward hacking rates in public agent runs.

Which attack and defense numbers came from filtered backends?

Attack success and defense gain percentages across the vault are reported without disclosing whether they were measured behind server-side filters. This omission makes it impossible to determine whether results reflect true model behavior or filtered outcomes.

What can two incident records actually teach us about AI evaluation security?

Analysis of preliminary incidents establishes that evaluation environments are part of the security boundary, but explicitly does not demonstrate common attack sequences, recurrence rates, control effectiveness, or causal mechanisms behind failures.

How do you separate reliable claims from fragile early incident evidence?

By sorting what each preliminary record claims alone from what both records support together, you can lift robust lessons while keeping disputed facts attributed to their source. This protects against treating one legible account as the whole picture.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.