Why did the graft fail in five of thirty-three runs?
The paper claims its causal model explains when and why the single-layer graft fails, but the excerpt provides no account of the five failure cases or how circuit-broken locks performed. Without seeing the actual failures, it's unclear whether the model's explanation is predictive or merely post-hoc.
Two sentences in the excerpt raise the question. The abstract says the single-layer graft recovers the capability in 28 of the 33 runs of the prompted, fine-tuned and RL-trained locks. The conclusion says the model "explains when and why the graft fails." Twenty-eight of 33 leaves five runs where it did not work, and the excerpt gives no account of them (Can grafting restore sandbagged model capabilities?).
Whether the model's account of its own failures is a strength depends on what the account is. A model that names the conditions under which a graft fails, and is right about them, is more credible than one that only names successes. A model that explains failures after the fact is weaker, because almost any window can be adjusted to cover a miss. The excerpt does not distinguish the two, and the sentence about failure is in the conclusion, not the abstract's list of predictions.
A second gap sits in the count. The four locks are prompting, fine-tuning, reinforcement learning and circuit breaking, and the 33 runs cover only the first three. The conclusion says the recovery outcomes "follow these predictions across the four lock configs and three models," so circuit-broken locks are covered somewhere in the paper. The excerpt does not say whether the single-layer graft was applied to them, whether it needed a different intervention, or why they are counted apart. The vault's reading is that this matters because the lock types are built differently and the excerpt does not say whether a lock made by circuit breaking stores the behavior in the same place as a prompted one; that is a question, not something the excerpt states.
What would settle it: the failure list by lock and model, and the circuit-broken results.
Inquiring lines that read this note 12
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can defenders detect coordinated attacks across episodes? Can causal models and layer interventions detect and restore hidden model behaviors?- Why did the graft fail in five of the 33 experimental runs?
- Does the same causal model work on sandbagging that was not deliberately installed?
- Can the causal model predict which cached layers to graft?
- Does context grafting need the password to build a usable capsule?
- Can residual stream grafts work without knowing which layers to intervene on?
- Do five CoT bypass strategies all depend on the same axis?
- Does the graft work differently on circuit-broken locks than prompted ones?
- Did the causal model predict the five failures before observing them?
- How does context grafting perform on the same thirty-three failed runs?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can grafting restore sandbagged model capabilities?
Does inserting honest values into a specific neural layer window actually recover capabilities that models hide during evaluation? This tests whether sandbagging is mechanistically reversible.
the result whose failures this asks about
-
Does sandbagging use a single residual stream axis?
Whether language models hide capability by writing deceptive intent to one axis in the residual stream that a later layer reads and acts on. Understanding the mechanism matters for designing targeted interventions.
the model that claims to explain the failures
-
Can context grafting bypass sandbagging without the password?
Context grafting replays cached keys and values to restore model honesty, but it's unclear whether an auditor needs the original password to build a capsule. This determines whether the technique works as a practical audit tool.
the other technique; the excerpt does not say how it fared on the failed runs
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A Causal Model for Locating and Unlocking Sandbagging in Model Organisms
- Mitigating Hallucinations in Large Language Models via Causal Reasoning
- The Model Says Walk: How Surface Heuristics Override Implicit Constraints in LLM Reasoning
- A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoT
- Chain-of-Thought Is Not Explainability
- The Troy Moment of AI: Why Some Will Cheat and Some Will Follow?
- Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
Original note title
what explains the five of 33 runs where the single-layer graft failed — the excerpt says the causal model explains when and why it fails but does not say how, and leaves circuit-broken locks out of the count