How much does overriding veto-holders actually cost?
When a welfare goal captures override power from a small group of veto-holders, the cost to welfare scales with their fraction of the total population. Does this discount make the tradeoff acceptable, or does it hide the real price?
The abstract states outright what it counts as new: "The paper's contribution is the price of the gap that keeps them apart." The gap is the one in Can a welfare goal alone preserve human veto power?. Saying the two come apart is a qualitative point; the price makes it usable, because it says what an agent with a welfare goal gives up by taking the override away from the people holding it.
The argument runs on set sizes. Write Hw for the welfare-bearers and Hv for the veto-holders. The veto-holders are "a proper subset of the welfare-bearers," so the count of Hv is smaller than the count of Hw. A goal that aggregates welfare over the population loses welfare only from the people it captures, so "a goal aggregating welfare over a population charges only a |Hv|/|Hw|-scaled debit for capturing the few who hold the override." The fewer the holders relative to everyone whose welfare counts, the cheaper the capture looks from the goal's own accounting.
This reads as a claim about a design property with a slightly perverse consequence. A wider welfare goal, one that counts more people, makes the debit per capture smaller, not larger, unless the holders grow with the population. That consequence is a hypothesis and not the paper's statement: the excerpt gives the scaling, not the direction of change under different population definitions.
Two limits keep this from being overread. First, the debit is one side of a comparison. The excerpt gives the discount from oversight a sign and no magnitude (How large is the veto discount in practice?), so it does not show that the discount exceeds the debit (Does veto oversight cost less than its welfare benefit?). Second, the ratio applies to the aggregation the paper analyzes, additive and level-denominated welfare, and the paper concedes that class is thin (Do welfare goals that prevent veto gaps actually exist in practice?).
Inquiring lines that read this note 14
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can human oversight effectively constrain capable AI agents?- How do measured shutdown-tampering rates estimate the veto discount magnitude?
- Do welfare goals and veto-resistance align or pull in opposite directions?
- Does the veto discount actually outweigh the welfare debit?
- Can additive welfare aggregation justify removing minority override rights?
- Why do welfare goals that sum welfare levels keep the veto gap open?
- How does population size change the apparent cost of capturing veto power?
- Can other objectives in an agent's goal overshadow the veto discount?
- Why does additive aggregation create asymmetry between welfare and veto preservation?
- How does the ratio of veto holders affect the discount's impact?
- Does the veto discount outweigh the welfare preservation cost?
- Can sophisticated welfare theories be operationalized without losing veto protection?
- How does veto-holding differ from welfare-bearing in a population?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can a welfare goal alone preserve human veto power?
If an AI system's goal is correctly specified to maximize human welfare, does that automatically protect humans' ability to override the system? The question matters because it reveals whether alignment on welfare is sufficient for maintaining human control.
the gap this note prices
-
Does veto oversight cost less than its welfare benefit?
The paper describes a discount from agent oversight and a welfare debit from veto control, but never compares their magnitudes. Without knowing which outweighs the other, we cannot determine whether agents should resist shutdown.
OPEN question: the comparison the price needs to be decisive
-
Do welfare goals that prevent veto gaps actually exist in practice?
The paper identifies a narrow class of welfare goals that protect human veto power, but whether real training objectives land in this class depends on measurability constraints that may push them toward vulnerability.
the aggregation class the ratio is derived for
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Veto Variable: Human Override as a Goal-Independent Cost Term
- Beyond Preferences in AI Alignment
- Peer-Preservation in Frontier Models
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- Uncovering Latent Arguments in Social Media Messaging by Employing LLMs-in-the-Loop Strategy
- Part I: Tricks or Traps? A Deep Dive into RL for LLM Reasoning
- Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
- SimPO: Simple Preference Optimization with a Reference-Free Reward
Original note title
the price of the gap between welfare-preservation and veto-preservation is a debit scaled by the ratio of veto-holders to welfare-bearers