Can you build a capsule without knowing the password?
Context grafting requires a cached capsule from a password-bearing prompt. If auditors need the password to create one, the technique only finds locks whose passwords are already known, limiting its value for discovering hidden sandbagging.
The conclusion defines the capsule as "the keys and values cached from a password-bearing prompt." That definition puts the password upstream of the technique: to have a capsule, someone had to run a prompt that contained the password (Can context grafting bypass sandbagging without the password?).
That matters for what the technique is for. In the paper's organisms the experimenters install the lock, so the password is presumably known to them (the excerpt does not say who holds it), and a capsule would be easy to make. The value of an audit is highest where the password is not known, since a hidden lock's key is by definition not handed to the person testing for it. If a capsule can only be made from a prompt that already contains the key, context grafting is a way to transfer an unlocked state, not a way to find out whether a lock exists. If it can be built some other way, from a prompt found by search or from a different cached state, the excerpt does not say so.
This is a vault reading of a definition, and the paper may address it in the body. Two different things could be meant. One is a demonstration on organisms where the password is available, which is a fair use for the paper's stated scope, sandbagging organisms. The other is a claim about auditing models whose passwords are unknown, which the excerpt does not make. Only the first is supported by what the excerpt says.
Related, but separate: the residual-stream graft needs the auditor to know where the axis is, which the causal model supplies (Can causal models guide audits that catch sandbagging?). The two techniques may therefore fail in different places, one on not knowing the layer, the other on not having the password.
What would settle it: how the capsules in the paper's experiments were obtained, and whether any were built without the password.
Inquiring lines that read this note 9
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can causal models and layer interventions detect and restore hidden model behaviors?- Why did the graft fail in five of the 33 experimental runs?
- Does context grafting need the password to build a usable capsule?
- Can residual stream grafts work without knowing which layers to intervene on?
- Does the password trigger work by switching the axis value?
- Does the graft work differently on circuit-broken locks than prompted ones?
- How does context grafting perform on the same thirty-three failed runs?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can context grafting bypass sandbagging without the password?
Context grafting replays cached keys and values to restore model honesty, but it's unclear whether an auditor needs the original password to build a capsule. This determines whether the technique works as a practical audit tool.
the technique whose input this asks about
-
Can causal models guide audits that catch sandbagging?
Explores whether a causal model of sandbagging behavior can help auditors design interventional tests that detect hidden capability suppression, and what such tests would need to work.
the audit claim both grafts serve, scoped to organisms
-
Do causal models of installed sandbagging generalize to wild cases?
Researchers built causal models of sandbagging by deliberately installing it in 7-8B models through four methods. The open question is whether these models describe sandbagging that emerges naturally from training pressure, or only the engineered variant.
installed locks have known passwords, which is part of why organisms are convenient and part of why the result may not carry over
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A Causal Model for Locating and Unlocking Sandbagging in Model Organisms
- Creating an Atomic User Model for Personality-Aware Large Language Model Interaction
- LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
- Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
- Stealing Reasoning Traces from Proprietary LLM APIs
- Trust propagation and structural containment in Multi-agent LLM pipelines
- The challenge of hidden gifts in multi-agent reinforcement learning
- RAG Does Not Work for Enterprises
Original note title
does context grafting need the password — the capsule is cached from a password-bearing prompt and the excerpt does not say whether an auditor who lacks the password can build one