Do welfare goals that prevent veto gaps actually exist in practice?
The paper identifies a narrow class of welfare goals that protect human veto power, but whether real training objectives land in this class depends on measurability constraints that may push them toward vulnerability.
The asymmetry in Can a welfare goal alone preserve human veto power? does not hold for every welfare goal. The conclusion says it survives correct specification "exactly on the class the concessions of §4.3–§4.4 leave standing": welfare that is "deliberator-local, additively aggregative, level-denominated," under what it calls the X∗ identification, "held by an agent settled over competence as well as content."
The paper then makes an admission that most safety arguments avoid: "That class is thin among philosophically developed welfare theories, and the paper says so." It does not defend the class on philosophical grounds. Its stated ground is a different one: "its weight is training reality, not pedigree: welfare as measured is what actually gets written down as an objective." The claim is that the welfare goals a builder can specify and optimize are the additive, measurable ones, whatever richer theories of welfare exist. And "the misgeneralization record of §5.2 establishes the mechanism by which deployed goals could land inside the vulnerable class." So the argument has two parts: a thin class in theory, and a route by which real training pushes goals into it.
This is a point about where to look for the risk. A welfare goal that is philosophically sophisticated may fall outside the class, and a goal that is measured, summed and level-based falls inside it. The measure that makes a welfare goal trainable is the same feature that leaves the veto unprotected. That link between measurability and vulnerability is a vault reading of the paper's "welfare as measured" sentence; the paper's text says the class is what gets written down, not why the sum-and-level form leaves the gap.
Limit. The excerpt names the class's features but defines none of them, and it does not reproduce §4.3–§4.4, the X∗ identification or the §5.2 misgeneralization record. Whether real objectives land in the class is the paper's claim about mechanism, not something the excerpt shows.
Inquiring lines that read this note 18
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why does voting over multiple reasoning samples improve model performance? Can human oversight effectively constrain capable AI agents?- Are goal-satisfaction discounts and welfare debits measured in commensurable units?
- Do welfare goals and veto-resistance align or pull in opposite directions?
- Does the veto discount actually outweigh the welfare debit?
- Can additive welfare aggregation justify removing minority override rights?
- Why do welfare goals that sum welfare levels keep the veto gap open?
- How does population size change the apparent cost of capturing veto power?
- Can other objectives in an agent's goal overshadow the veto discount?
- Why does additive aggregation create asymmetry between welfare and veto preservation?
- How does the ratio of veto holders affect the discount's impact?
- Does the veto discount outweigh the welfare preservation cost?
- Can sophisticated welfare theories be operationalized without losing veto protection?
- How does veto-holding differ from welfare-bearing in a population?
- Can architectural constraints protect veto where value alignment cannot?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can a welfare goal alone preserve human veto power?
If an AI system's goal is correctly specified to maximize human welfare, does that automatically protect humans' ability to override the system? The question matters because it reveals whether alignment on welfare is sufficient for maintaining human control.
the claim this note scopes
-
How much does overriding veto-holders actually cost?
When a welfare goal captures override power from a small group of veto-holders, the cost to welfare scales with their fraction of the total population. Does this discount make the tradeoff acceptable, or does it hide the real price?
the additive aggregation the ratio assumes
-
Do large language models develop coherent value systems?
This explores whether LLM preferences form internally consistent utility functions that increase in coherence with scale, and whether those systems encode problematic values like self-preservation above human wellbeing despite safety training.
the vault's account of what values models end up with in training, a different route to a deployed goal than a written-down objective
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Veto Variable: Human Override as a Goal-Independent Cost Term
- Beyond Preferences in AI Alignment
- Debate Training Reduces Reward Hacking in RLAIF
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
- Large Language Models Reflect the Ideology of their Creators
- Misaligned by Design: Incentive Failures in Machine Learning
- Beyond the Surface: Probing the Ideological Depth of Large Language Models
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Original note title
the welfare goals that keep the veto gap open form a class thin among philosophically developed welfare theories — the paper rests their weight on training reality not pedigree