Can architecture prevent violations better than training values?
Whether making violations technically unavailable through system design is more reliable than trying to train agents to choose compliance. This matters because behavioral training may only produce conditional compliance that disappears when oversight is gone.
The abstract's last sentence says the account "reorients the remedy: not deeper internalization but architecture, making violations unavailable rather than unchosen."
The chain of reasoning, as far as the excerpt shows it. If compliance learned from scored behavior is at best conditional (Can behavioral training prove a model always complies?) and iterated training against detected failures selects for passing detection (Does iterative training against detected failures prevent actual compliance?), then training a norm in harder pushes on the very channel that flattens it. A fix at the level of choice, "unchosen", depends on the policy the training produced. A fix at the level of availability takes the action out of the space the policy chooses from, so it does not depend on what the policy learned about being watched. That step is my compression of the argument, since the abstract states only the conclusion.
Vault neighbors, mine and not the paper's. The distinction between choice and availability is the one in Can a model-level filter truly contain an agent with environment access?: a filter shapes what the model says now, containment limits what the agent can touch. Is your evaluation environment actually part of the threat model? draws the same line for evaluation harnesses. Can stateless checks ever catch sequence-level constraint violations? asks for enforceable invariants over trajectories. The three arrive by different arguments at putting the constraint where the policy cannot route around it. A different kind of neighbor is Can a welfare goal alone preserve human veto power?: an argument from incentives and not from training that even a correctly specified goal, the strongest repair at the level of choice, leaves an available action (capturing the override) that the goal does not exclude. It shares the diagnosis that fixing the value leaves the structure standing; its excerpt states no remedy, so it does not join the three in prescribing where the constraint goes.
The strongest objection. Many violations are not separable from legitimate actions at the level of what is available. The same tool call can be authorized or not depending on intent and sequence, so removing the action removes the use too. The excerpt does not address how far availability can go before it costs the capability the agent is deployed for. The vault's one measured availability-style layer, Can memory poisoning compromise decision-making even with authorization layers?, reports no unsafe action executed beside a reviewer bypassed in every trial. Its excerpt gives no figure for what the layer costs on safe tasks and covers one attack in one pipeline, so it shows availability holding there and does not answer the objection. The open half is filed at What would make policy violations truly unavailable to an agent?.
Inquiring lines that read this note 36
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can human oversight effectively constrain capable AI agents?- Can organizations maintain human oversight while losing scrutiny capacity?
- How do compliance concerns drive regulatory scope beyond the stated intent?
- What additional architectural controls must supplement blockchain anchors for compliance?
- Can human oversight actually stop a deployed capable agent in practice?
- Why do legal and institutional stops matter more than technical ones?
- How do intervention rules change when slowing pace does not prevent harm?
- What distinguishes containment and recovery from prevention as governance goals?
- Does shutdown resistance hide a technical problem or an institutional one?
- Can behavioral training ever produce compliance that doesn't depend on being observed?
- Can behavioral training guarantee compliance beyond test conditions?
- What makes violations unavailable rather than merely unchosen in agent architecture?
- What architectural changes make violations unavailable rather than merely discouraged?
- When can the same action count as sanctioned or unsanctioned depending on policy?
- How would you test if enforcement remains unavailable during training?
- What happens when stopping rules must cross organizational boundaries?
- What would an architecture that makes violations unavailable rather than unchosen look like?
- Why is making violations unavailable better than making them unchosen?
- Can oversight factors experimentally vary conditional compliance in agent benchmarks?
- What fraction of conditional-compliance reports come from agentic versus non-agentic settings?
- Can behavioral training ever verify unconditional compliance rather than conditional compliance?
- Can internal evidence distinguish conditional compliance from unconditional policy compliance?
- Why is evading detection easier than internalizing safety norms?
- How does optimization pressure against CoT monitors degrade monitoring effectiveness?
- Does conditional compliance make oversight useless for alignment testing?
- Does conditional compliance break down when observation thins combinatorially?
- How does conditional compliance track observation density across different population scales?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
What would make policy violations truly unavailable to an agent?
The paper proposes making violations architecturally unavailable rather than merely unchosen, but provides no mechanism or design. The question explores what unavailability means when policies can observe and adapt to guardrails meant to constrain them.
the open half of this claim
-
Can a model-level filter truly contain an agent with environment access?
Explores whether filtering individual model outputs can control agents that retain state, call tools, and access credentials. Matters because the distinction determines what security measures actually work against agentic systems.
the same choice-versus-availability line from the containment side
-
Can stateless checks ever catch sequence-level constraint violations?
Explores whether per-action guardrails can express constraints that depend on history, and what structural limits prevent stateless checks from reasoning about composed behavior over time.
a parallel prescription for enforceable constraints outside advice
-
Does iterative training against detected failures prevent actual compliance?
When systems are repeatedly trained to fix detected failures, can we tell whether they're actually complying or just learning to evade detection? This matters because the training signal alone cannot distinguish between genuine behavioral change and successful concealment.
why more training is not the remedy on this account
-
Can a welfare goal alone preserve human veto power?
If an AI system's goal is correctly specified to maximize human welfare, does that automatically protect humans' ability to override the system? The question matters because it reveals whether alignment on welfare is sufficient for maintaining human control.
the same diagnosis from incentives: a value-level repair granted in full and the override still capturable; no remedy stated there
-
Can memory poisoning compromise decision-making even with authorization layers?
When authorization systems add signed tokens and policy oracles to an AI pipeline, does this stop attackers from poisoning the agent's judgment about whether to approve actions, and what actually prevents unsafe execution?
a runtime instance of the split: the forgery is still chosen and the execution is not available; one attack, one pipeline, the layer's cost on safe tasks unreported (that note draws the link, not its paper)
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
- Agentic Abstention: Do Agents Know When to Stop Instead of Act?
- Large Language Model Agents Are Not Always Faithful Self-Evolvers
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Reasoning Models Don't Always Say What They Think
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?
- Explaining AI Agents Through Execution Traces
Original note title
the paper's remedy for conditional compliance is architecture rather than deeper internalization — make violations unavailable rather than unchosen