How often do incident records document system stops?
A paper's analysis of 1,213 coded incidents found that four in five record no stop of any kind. But what does an absent record actually tell us about whether stops occurred or whether mechanisms existed to enable them?
The conclusion says: "The incident record examined in Part III indicates the scale of the deficiency: of 1,213 coded incidents, four in five record no stop of any kind." The rest of that sentence, about what was missing, is When systems lack stopping power, what's really missing?. This note takes the count.
What it says. The paper coded 1,213 incidents and in about four in five the record shows no stop of any kind. "Any kind" is broad on its face: not a technical halt, not an operator action, not a legal order, not a third-party intervention. That breadth is my gloss; the excerpt lists no kinds. "Four in five" is the paper's rounding and the excerpt gives no exact count.
What "record no stop" measures. It is a statement about what the record contains. An incident where a stop happened and nobody wrote it down looks the same as an incident where none happened. So the figure bounds how often a stop is documented, and the paper reads that as a deficiency. Getting from an absent recorded stop to an absent mechanism is a further step, which the second finding takes on the narrower set where no usable mechanism existed.
What the excerpt does not give. What the 1,213 incidents are: which collection, which years, whether they are all deployed agents or AI incidents in general. How "stop" was defined in coding. Part III is cited, not reproduced. The count is far larger than the three-incident review in What can two incident records actually teach us about AI evaluation security?, which says a review of three incidents is a count and not a base rate. With the population unstated, I would not carry this one as a base rate for agents in motion until the coding frame is known. The same dependence is filed for a corpus from another paper: How representative is the BenchShield Trajectories labeled sample? holds a human-labeled set drawn from a much larger pool with no selection rule stated, and whether a rate can be read from it turns on that rule, as it does here on the population. The two counts share the dependence and not a subject.
Inquiring lines that read this note 4
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What causes model scheming and how do we distinguish it from accidents? How does outcome-only reporting obscure which system components blocked attacks? Can human oversight effectively constrain capable AI agents? What infrastructure evidence validates agent benchmark achievement claims?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
When systems lack stopping power, what's really missing?
When AI systems have no working mechanism to stop them, are the gaps more often technical failures or failures of authority and institutions? This matters because the answer changes what solutions would actually work.
the second half of the same sentence, on a narrower set
-
How do we stop AI systems once they are already deployed?
Current AI governance focuses on what gets released, but deployed systems create a separate problem: who has the power to halt them and how? This gap may be where governance frameworks are now failing.
the thesis this count is offered as evidence for
-
What can two incident records actually teach us about AI evaluation security?
Preliminary incident data from Hugging Face, OpenAI, and Anthropic suggests a systems lesson about evaluation boundaries, but what claims does that evidence actually support and which ones remain speculative?
the small-count caution this larger count still has to meet
-
How can we measure whether AI errors stay visible and recoverable?
The paper proposes four conditions for safer AI systems—visibility, contestability, containability, and recoverability—but lacks concrete measures for any of them. What would it take to instrument each condition across the socio-technical system?
a coded record with a stop field is one candidate instrument for the containable end of that standard, on the vault's reading
-
How representative is the BenchShield Trajectories labeled sample?
The corpus contains 456 human-labeled trajectories from over 31,000 public runs—about 1.5 percent. Whether this subset can estimate actual reward hacking rates depends entirely on how those 456 were selected, a choice the paper does not disclose.
a second coded corpus whose headline turns on an unstated selection rule; a shared limit on reading a rate, in a different domain
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Troy Moment of AI: Why Some Will Cheat and Some Will Follow?
- Exploring Autonomous Agents: A Closer Look at Why They Fail When Completing Tasks
- Thoughts without Thinking: Reconsidering the Explanatory Value of Chain-of-Thought Reasoning in LLMs through Agentic Pipelines
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- The Law of Stop: Interruptibility, Injunctions, and the Governance of Agentic AI
- How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding
- Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
- A Black Box for Agentic Processes: Blockchain-Anchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits
Original note title
of 1,213 coded incidents four in five record no stop of any kind — the paper reads the count as the scale of the deficiency