SYNTHESIS NOTE
Topics›Alignment›this note

Why do safety failures remain invisible to our evaluation methods?

Current evaluation practices assume failures are obvious, localized, and immediate. But as AI systems deploy into workflows, failures are becoming quiet, distributed, and normalized before detection. What blindspots does this mismatch create?

Synthesis note · 2026-09-23 · sourced from Alignment

The word "hidden" in the title invites a misreading, and the paper closes it off: the challenges "are not hidden because they are mystical or technically invisible. They are hidden because our dominant habits of evaluation, interface design, and governance still assume that safety failures are mostly local, output-level, and immediately legible." The blindness is inherited from instruments built for a different failure shape.

The failure shape it says has changed is stated in the abstract: in deployed systems, many of the most consequential failures are "plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards." Each clause breaks one of the three assumptions. Plausible failures are not immediately legible. Distributed failures are not local. And a failure spread over components has no single output to inspect. A shocking output, a policy-violating generation or a vivid adversarial example are what the discourse is "well prepared to notice"; the paper says it is "less prepared" for the quiet forms once models are embedded in ordinary work.

This makes the problem an instrumentation problem, not a research-frontier problem. If the failures are visible in principle and unmeasured in practice, then the fix is to build the instrument, which is what the title's "not instrumenting" says, and the open question that follows, what an instrument for the paper's own standard would measure, is How can we measure whether AI errors stay visible and recoverable?. It also explains why a green benchmark is weak evidence. A benchmark is an output-level, local, snapshot instrument, so it can be clean exactly where the paper says the hazard sits; the snapshot case is Can safety tests miss hazards that build over time?.

Two notes from other papers give the local and output-level assumptions a concrete form. Can individual components pass safety checks if the system still fails? is the not-local case: the check available per component tests a different property from the one that matters for the system. Can a correct outcome hide protocol violations in multi-agent systems? is the output-level case: the outcome passes whether or not the required step ran. The first is an argument across three excerpts and the second a single-setting result, and pairing either with this paper's diagnosis is the vault's reading.

What the excerpt does not give. No case study, dataset or incident is cited for any of the three assumptions or for the claim that failures have changed shape, so this is a diagnosis by argument. The vault holds independent evidence for parts of it, linked below, but that evidence comes from other papers and was not brought together by this one.

Inquiring lines that read this note 26

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How does outcome-only reporting obscure which system components blocked attacks? How can workflow-level validation detect semantic corruption that protocol compliance misses? How do coordinated agent sequences violate constraints that individual actions respect? What determines whether AI system errors remain visible and contestable? Do AI capability benchmarks accurately measure reasoning ability or just surface patterns? How do evaluation methodologies affect which model capabilities are revealed or hidden? How does training data contamination persist through safety alignment mechanisms?

Related concepts in this collection 10

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
19 direct connections · 173 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

hidden safety-critical challenges are hidden by habits of evaluation, interface design, and governance that assume failures are local, output-level, and immediately legible — not by technical invisibility