SYNTHESIS NOTE
Topics›Alignment›this note

How much does overriding veto-holders actually cost?

When a welfare goal captures override power from a small group of veto-holders, the cost to welfare scales with their fraction of the total population. Does this discount make the tradeoff acceptable, or does it hide the real price?

Synthesis note · 2026-09-23 · sourced from Alignment

The abstract states outright what it counts as new: "The paper's contribution is the price of the gap that keeps them apart." The gap is the one in Can a welfare goal alone preserve human veto power?. Saying the two come apart is a qualitative point; the price makes it usable, because it says what an agent with a welfare goal gives up by taking the override away from the people holding it.

The argument runs on set sizes. Write Hw for the welfare-bearers and Hv for the veto-holders. The veto-holders are "a proper subset of the welfare-bearers," so the count of Hv is smaller than the count of Hw. A goal that aggregates welfare over the population loses welfare only from the people it captures, so "a goal aggregating welfare over a population charges only a |Hv|/|Hw|-scaled debit for capturing the few who hold the override." The fewer the holders relative to everyone whose welfare counts, the cheaper the capture looks from the goal's own accounting.

This reads as a claim about a design property with a slightly perverse consequence. A wider welfare goal, one that counts more people, makes the debit per capture smaller, not larger, unless the holders grow with the population. That consequence is a hypothesis and not the paper's statement: the excerpt gives the scaling, not the direction of change under different population definitions.

Two limits keep this from being overread. First, the debit is one side of a comparison. The excerpt gives the discount from oversight a sign and no magnitude (How large is the veto discount in practice?), so it does not show that the discount exceeds the debit (Does veto oversight cost less than its welfare benefit?). Second, the ratio applies to the aggregation the paper analyzes, additive and level-denominated welfare, and the paper concedes that class is thin (Do welfare goals that prevent veto gaps actually exist in practice?).

Inquiring lines that read this note 14

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Can human oversight effectively constrain capable AI agents? Why does voting over multiple reasoning samples improve model performance? Can aggregate reward models represent diverse human preferences without bias?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 71 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

the price of the gap between welfare-preservation and veto-preservation is a debit scaled by the ratio of veto-holders to welfare-bearers