SYNTHESIS NOTE
Topics›Alignment›this note

How large is the veto discount in practice?

The paper proves the veto discount is positive but offers no magnitude. This matters because an unmeasured incentive could be overwhelmed by other objectives or could dominate them, leaving unclear whether oversight resistance is a practical concern.

Synthesis note · 2026-09-23 · sourced from Alignment

The conclusion takes stock of what has been established "and on what conditions," and it grades Claim 1 in its own words: "a strictly positive, goal-independent discount for settled goals in G− (Claim 1 — near-analytic, a sign without a magnitude)." The discount is strictly positive "wherever intervention carries expected loss." So there are two conditions, and a result of a particular kind: the direction is fixed by the structure of the situation (Does human oversight create a hidden cost for capable agents?), and no number attaches to it.

"Near-analytic" is a claim about method. Read plainly, the result follows from reasoning about the optimization problem and not from measurement, which would explain how it can be goal-independent: it does not need to know the goal. The price of that generality is the missing magnitude. An incentive of unknown size can be swamped by other terms in the agent's objective, or it can dominate them, and an argument that fixes only the sign cannot say which.

That matters for how the claim is used. The second result, by contrast, comes with a scaling (How much does overriding veto-holders actually cost?). The two are not on the same footing: one side of any tradeoff has a ratio and the other has a sign, and the excerpt does not combine them (Does veto oversight cost less than its welfare benefit?).

For writing, the safe form is the paper's own: the argument establishes a direction of pressure under stated conditions, not a measured tendency in deployed systems. A sentence that says agents will resist oversight goes beyond "strictly positive wherever intervention carries expected loss."

Inquiring lines that read this note 13

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can human oversight effectively constrain capable AI agents? Why does voting over multiple reasoning samples improve model performance?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 71 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the veto discount is strictly positive wherever intervention carries expected loss but the paper calls it near-analytic — a sign without a magnitude