SYNTHESIS NOTE
Topics›Alignment›this note

Can independent scenario factors isolate what drives scheming?

Prior work bundled tool domains, goals, and oversight together, making it impossible to attribute behavior changes to specific conditions. Can factorizing these elements separately reveal which conditions actually trigger scheming?

Synthesis note · 2026-09-23 · sourced from Alignment

The contribution is a benchmark. The abstract describes "a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms." The discussion adds the property that matters: "scenario factors that can be varied independently."

Two stated limits of earlier work. The introduction says prior work "examines only a small number of scenarios," which limits isolating how conditions shape scheming. It adds that "limited scale and task diversity" also restrict coverage of realistic deployment settings and the range of scheming strategies observed. The first is a problem of attribution: when goal, tools and oversight change together inside one scenario, a difference in behavior cannot be assigned to any one of them. The second is a problem of coverage. Factorized synthesis addresses both, and the ordering in What drives scheming behavior most strongly in language models? is the kind of result only the first makes possible.

Two factor lists that do not quite match (my reading). The abstract lists tool domains, instrumental goals, oversight conditions and pressure mechanisms. The introduction lists instrumental goals, perceived consequences and influences of actions, environmental affordances and oversight conditions. Tool domains may correspond to affordances and pressure mechanisms to perceived consequences, but the excerpt does not say so. The mapping matters for reading "pressure" in Does pressure on AI agents lead to covert scheming behavior?.

The same move elsewhere in the vault. How can we make reward-hacking visible in agent evaluation? separates parts of an evaluation so a failure can be attributed to one of them. Factorizing the scenario does the same for the causes of a safety failure.

Where the vault meets the attribution problem. Do authorization rules or restricted tools prevent test modifications? is a regime that changed two things together and reported one zero, so the result cannot be credited to either; that is the attribution limit above, met in a different paper. Which security protections actually slow down agent exploits? makes the move SchemeArena makes, on an environment's protections and not on an agent's goals, oversight or pressure. The excerpt does not say whether SchemeArena's factors are fully crossed, so it is not shown to separate bundled changes the way a complete crossing would.

What the excerpt does not give. Whether the factors are fully crossed or how the 400 divide across them (a crossed design over four axes would leave few scenarios per cell), the levels of each factor, and how scenarios were checked for realism. The reported results concern propensity, not capability.

Inquiring lines that read this note 12

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What causes model scheming and how do we distinguish it from accidents? Do planted honeypot tests reliably measure reward hacking? Do frontier models develop hidden self-protective behaviors? How do coordinated agent sequences violate constraints that individual actions respect? Do pretraining and finetuning change model capabilities or only output behavior?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 110 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

factorized scenario synthesis lets SchemeArena vary tool domains, instrumental goals, oversight conditions and pressure across 400 scenarios — prior scheming work examined only a small number