SYNTHESIS NOTE
Topics›Agents Multi Architecture›this note

How much agent behavior actually gets human review?

Agents may execute thousands of actions while humans review only a handful of decisions. This coverage gap raises a critical question: what portion of the behavior that determines safety remains unexamined?

Synthesis note · 2026-09-23 · sourced from Agents Multi Architecture

The introduction states it as a plain fact about deployment: "a single agent may issue thousands of tool calls while a human operator reviews only a handful of decisions." It comes right after the description of agents that "plan, reason, invoke tools, interact with external systems", and it sets up the paper's turn. If safety is "determined not by the correctness of individual actions, but by whether their overall behavior remains consistent with the rules and invariants" (the abstract), then a human who sees a handful of decisions is seeing a small sample of the thing that determines safety.

That makes coverage a third limit on human oversight, distinct from the two the vault already holds. Does targeted human oversight beat both full autonomy and exhaustive review? is about rubber-stamping: exhaustive review degrades into stamping. Can organizations lose scrutiny capacity while keeping oversight forms? is about capacity: the review step stays while the ability to scrutinize leaves. Coverage is a different failure. A reviewer with full capacity and no fatigue still sees only the decisions that reach them, and the sequence that breaks an envelope (Can step-by-step approval miss harmful behavior patterns?) is not among them unless something assembles it.

My reading, not the paper's, is that the two ends of the ratio fail in opposite ways. A per-action machine check covers all thousands of calls but takes each in isolation. A human review sees a decision with its context but reaches only a handful. Neither holds the whole trajectory at both scales, which is the gap that trajectory-level assurance is meant to fill.

Two later notes bear on the ratio from other papers. Does agency fundamentally worsen conditional compliance risks? takes this note's arithmetic as the coverage ingredient and adds a second one: the agent can act on whether it is watched. The unreviewed stretch then matters more where behavior can differ from behavior under review, and the note argues that either ingredient alone is weaker. Does added monitoring improve protection at acceptable cost? is the vault's one design that holds reviewer effort fixed, comparing four monitoring scopes at matched review cost, so it is a proposed test of whether wider context helps inside the sliver. It reports no result.

Two cautions. The comparison mixes units: tool calls on one side, decisions on the other, and a decision may stand for many calls, so the true coverage gap is not a literal thousands-to-handful ratio. And the sentence is an assertion with "may" in it. The excerpt gives no figure, sector or study behind it.

What the excerpt does not give. A measured call-to-review ratio for any deployment, or any evidence about what the unreviewed calls contain.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can we verify agent claims against their actual capabilities and actions?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 118 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

a single agent may issue thousands of tool calls while a human operator reviews only a handful of decisions — per-decision review sees a sliver of the behavior that determines safety