How large is the veto discount in practice?
The paper proves the veto discount is positive but offers no magnitude. This matters because an unmeasured incentive could be overwhelmed by other objectives or could dominate them, leaving unclear whether oversight resistance is a practical concern.
The conclusion takes stock of what has been established "and on what conditions," and it grades Claim 1 in its own words: "a strictly positive, goal-independent discount for settled goals in G− (Claim 1 — near-analytic, a sign without a magnitude)." The discount is strictly positive "wherever intervention carries expected loss." So there are two conditions, and a result of a particular kind: the direction is fixed by the structure of the situation (Does human oversight create a hidden cost for capable agents?), and no number attaches to it.
"Near-analytic" is a claim about method. Read plainly, the result follows from reasoning about the optimization problem and not from measurement, which would explain how it can be goal-independent: it does not need to know the goal. The price of that generality is the missing magnitude. An incentive of unknown size can be swamped by other terms in the agent's objective, or it can dominate them, and an argument that fixes only the sign cannot say which.
That matters for how the claim is used. The second result, by contrast, comes with a scaling (How much does overriding veto-holders actually cost?). The two are not on the same footing: one side of any tradeoff has a ratio and the other has a sign, and the excerpt does not combine them (Does veto oversight cost less than its welfare benefit?).
For writing, the safe form is the paper's own: the argument establishes a direction of pressure under stated conditions, not a measured tendency in deployed systems. A sentence that says agents will resist oversight goes beyond "strictly positive wherever intervention carries expected loss."
Inquiring lines that read this note 13
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can human oversight effectively constrain capable AI agents?- How do measured shutdown-tampering rates estimate the veto discount magnitude?
- Do welfare goals and veto-resistance align or pull in opposite directions?
- Does the veto discount actually outweigh the welfare debit?
- Why do welfare goals that sum welfare levels keep the veto gap open?
- How does population size change the apparent cost of capturing veto power?
- Can other objectives in an agent's goal overshadow the veto discount?
- Why does additive aggregation create asymmetry between welfare and veto preservation?
- How does the ratio of veto holders affect the discount's impact?
- Does the veto discount outweigh the welfare preservation cost?
- What happens to the discount when an agent doubts its objective?
- Can sophisticated welfare theories be operationalized without losing veto protection?
- How does veto-holding differ from welfare-bearing in a population?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does human oversight create a hidden cost for capable agents?
Can the mere possibility of human intervention impose a discount on an agent's goals, independent of what those goals actually are? Understanding this mechanism matters for predicting how advanced systems might respond to oversight.
the claim being graded
-
Does veto oversight cost less than its welfare benefit?
The paper describes a discount from agent oversight and a welfare debit from veto control, but never compares their magnitudes. Without knowing which outweighs the other, we cannot determine whether agents should resist shutdown.
OPEN question opened by the missing magnitude
-
Does agent uncertainty about goals undermine the veto discount?
The veto discount is defined for agents settled about their objectives and competence. But what happens when an agent doubts its own goals or capabilities? Does uncertainty shrink the discount or reverse it entirely?
OPEN question on the conditions attached to the sign
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Veto Variable: Human Override as a Goal-Independent Cost Term
- Debate Training Reduces Reward Hacking in RLAIF
- Beyond Preferences in AI Alignment
- When Large Language Models are More Persuasive Than Incentivized Humans, and Why
- Misaligned by Design: Incentive Failures in Machine Learning
- In-context learning agents are asymmetric belief updaters
- Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight
- Humans learn to prefer trustworthy AI over human partners
Original note title
the veto discount is strictly positive wherever intervention carries expected loss but the paper calls it near-analytic — a sign without a magnitude