SYNTHESIS NOTE
Topics›Alignment›this note

Why did the graft fail in five of thirty-three runs?

The paper claims its causal model explains when and why the single-layer graft fails, but the excerpt provides no account of the five failure cases or how circuit-broken locks performed. Without seeing the actual failures, it's unclear whether the model's explanation is predictive or merely post-hoc.

Synthesis note · 2026-09-23 · sourced from Alignment

Two sentences in the excerpt raise the question. The abstract says the single-layer graft recovers the capability in 28 of the 33 runs of the prompted, fine-tuned and RL-trained locks. The conclusion says the model "explains when and why the graft fails." Twenty-eight of 33 leaves five runs where it did not work, and the excerpt gives no account of them (Can grafting restore sandbagged model capabilities?).

Whether the model's account of its own failures is a strength depends on what the account is. A model that names the conditions under which a graft fails, and is right about them, is more credible than one that only names successes. A model that explains failures after the fact is weaker, because almost any window can be adjusted to cover a miss. The excerpt does not distinguish the two, and the sentence about failure is in the conclusion, not the abstract's list of predictions.

A second gap sits in the count. The four locks are prompting, fine-tuning, reinforcement learning and circuit breaking, and the 33 runs cover only the first three. The conclusion says the recovery outcomes "follow these predictions across the four lock configs and three models," so circuit-broken locks are covered somewhere in the paper. The excerpt does not say whether the single-layer graft was applied to them, whether it needed a different intervention, or why they are counted apart. The vault's reading is that this matters because the lock types are built differently and the excerpt does not say whether a lock made by circuit breaking stores the behavior in the same place as a prompted one; that is a question, not something the excerpt states.

What would settle it: the failure list by lock and model, and the circuit-broken results.

Inquiring lines that read this note 12

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can defenders detect coordinated attacks across episodes? Can causal models and layer interventions detect and restore hidden model behaviors? What limitations prevent automated research from matching human research quality? How do agents balance task completion with privacy compliance and security?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 88 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

what explains the five of 33 runs where the single-layer graft failed — the excerpt says the causal model explains when and why it fails but does not say how, and leaves circuit-broken locks out of the count