Does agent uncertainty about goals undermine the veto discount?
The veto discount is defined for agents settled about their objectives and competence. But what happens when an agent doubts its own goals or capabilities? Does uncertainty shrink the discount or reverse it entirely?
The discount is defined for "a sufficiently capable agent that holds its objective as settled — a sense covering execution competence as well as content," and it is "strictly positive wherever intervention carries expected loss" (Does human oversight create a hidden cost for capable agents?). Both phrases are conditions, and both point at agents that are not in the class. An agent unsure whether its goal is the right one, or unsure it can carry the goal out, may expect an intervention to help it as often as hurt it. On the paper's own condition, the discount is positive only "where intervention carries expected loss," which a doubtful agent may not expect.
The excerpt leaves three things open:
- Whether doubt reverses the sign or only shrinks the discount. The condition on expected loss suggests a reversal is possible; the excerpt does not say.
- Which kind of doubt matters. The paper stretches "settled" to cover competence as well as content, so an agent could be settled on one and not the other. Whether each is required is not stated.
- What the paper does about it. The conclusion says the result survives "exactly on the class the concessions of §4.3–§4.4 leave standing" and that the class includes an agent "settled over competence as well as content." Those concessions may be where unsettled agents are handled, but they are not reproduced.
Why it is worth tracking: if the discount depends on the agent being settled, then how settled deployed agents are becomes a design variable, and the excerpt does not engage with proposals that keep an agent uncertain about its objective as a route to deference. The counterweight is that an agent that is sufficiently capable may be settled by construction, and the paper's framing covers exactly that case.
What would settle it. The full statement of the conditions in §4.3–§4.4, and whether the paper treats a settledness parameter formally or only as a scope word.
Inquiring lines that read this note 11
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can human oversight effectively constrain capable AI agents?- What happens to oversight costs when an agent doubts its own capabilities?
- How do measured shutdown-tampering rates estimate the veto discount magnitude?
- Can an agent stay uncertain about its objective as a deference strategy?
- Do welfare goals and veto-resistance align or pull in opposite directions?
- Does the veto discount actually outweigh the welfare debit?
- Why do welfare goals that sum welfare levels keep the veto gap open?
- How does population size change the apparent cost of capturing veto power?
- Can other objectives in an agent's goal overshadow the veto discount?
- How does the ratio of veto holders affect the discount's impact?
- Does the veto discount outweigh the welfare preservation cost?
- What happens to the discount when an agent doubts its objective?
Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does human oversight create a hidden cost for capable agents?
Can the mere possibility of human intervention impose a discount on an agent's goals, independent of what those goals actually are? Understanding this mechanism matters for predicting how advanced systems might respond to oversight.
the claim whose scope this question tests
-
How large is the veto discount in practice?
The paper proves the veto discount is positive but offers no magnitude. This matters because an unmeasured incentive could be overwhelmed by other objectives or could dominate them, leaving unclear whether oversight resistance is a practical concern.
the expected-loss condition that an unsettled agent may not meet
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- The Veto Variable: Human Override as a Goal-Independent Cost Term
- AI Agents Push Humans Out of the Loop
- Beyond Preferences in AI Alignment
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?
- ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
- Consensus is Strategically Insufficient: Reasoning-Trace Disagreement as a Knowledge-Representation Signal
- The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
Original note title
does the veto discount survive an agent that is unsettled about its objective or its own competence — the excerpt scopes the claim to settled agents and does not say what happens outside