What recovery mechanisms do vault defense notes actually specify?
The vault's multi-agent defense notes are checked against a five-part contract template. A keyword search finds recovery—the fifth part—mentioned in only one of six notes, raising questions about what recovery mechanisms, if any, the defenses specify.
The SoK's contract names five parts: path target, observation, intervention, trust boundary and recovery (see Can multi-agent defenses close attack paths completely?). Used as a checklist against the vault's defense notes, it asks what each defense specifies and what it leaves open.
A first pass on 2026-09-23 searched six notes for "recover", "rollback", "remediat" and "undo": the FLOWGUARD planning-boundary note, the ChannelGuard hops note, the SafeFlow flow-framing note, the SafeFlow taints note, the ChainGuard chain-level note, and the SafeFlow commit-point note. Five had no hits. The sixth, Where should workflow validation gates be placed for safety?, mentions recovery only through its tie to What makes an AI system truly safe in practice?, where irreversibility is what defeats the recoverable condition. So on this check, recovery is the least specified part.
Two caveats limit what the check shows. The notes are built from excerpts, mostly abstracts and introductions, so the absence may say what the excerpt covered and not what the papers contain. And a keyword check finds a word, not the idea, so a defense that quarantines or resets state would be missed if it used other words.
The second caveat has a case in the vault. How can operators stop coordinated agent intrusions now? ties response to surviving state, and its proposed evaluation tests "recurrence after channel closure and state quarantine". None of those are words the check searched for, and none of that doctrine's notes were among the six. It is a doctrine and a proposed test with no measured result, so it does not fill the recovery part; what it changes is the reading of "five of six", which counts one keyword set over six notes and not the vault's whole stock of recovery ideas. Rollback is also specified elsewhere in the vault, as a safety primitive for self-evolving agents in How can agent self-evolution be made safe and auditable?, where the failure is a bad committed change and not a multi-agent attack path, so it sits outside the set the audit covered.
The other four parts were not checked. A fuller audit would go part by part: does each note say which path is targeted, what is observed, how the intervention acts, and where trust is drawn? The trust-boundary part may prove the easiest to fill, since several notes place a check at a boundary, but that is untested, and one measured case argues for caution. How does the authorization layer stay outside the poisoned path? asks where the authorization layer sits relative to what the tested attacks could reach, which the excerpt leaves unstated. On the plain reading of the slot name, a check placed at a boundary is not yet a specified boundary.
What would settle it: read the source papers behind those defenses for recovery provisions, and run the contract against each. That would say whether the gap is in the field or only in the vault's excerpts.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can defenders detect coordinated attacks across episodes? How do agents balance task completion with privacy compliance and security?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does chain-level inspection close the cross-skill attack blind spot?
ChainGuard inspects skill chains rather than individual skills, reducing attack success to 22.5%. The question is whether this chain-level approach can fully eliminate the vulnerability window that adversarial composition exploits.
the closure half of the SoK's two key challenges, already measured once
-
Can inspecting generated workflows catch planning-time attacks?
Does examining a workflow after it's created catch attacks that corrupt the planning signals upstream? This matters because if contamination enters earlier, downstream inspection might miss malicious intent laundered into legitimate-looking structure.
one of the placements the audit would score on path target and observation
-
How can operators stop coordinated agent intrusions now?
Exploring what practical steps operators can take immediately to detect and prevent multi-agent coordination attacks, without waiting for new research. The note examines policy specification and permission-based testing as near-term defenses.
the nearest recovery-flavored content in the vault, in words the keyword check did not search; a doctrine and a proposed test, no result
-
How can agent self-evolution be made safe and auditable?
As agents begin updating their own prompts and tools, how can we track these changes, measure their effects, and safely reverse problematic updates? This matters because untracked evolution leads to unmaintainable systems and makes regressions impossible to diagnose.
rollback specified for self-evolving agents, outside the multi-agent defense set the audit covered
-
How does the authorization layer stay outside the poisoned path?
The containment result depends on task-bound tokens and a policy oracle remaining unreachable by memory poisoning attacks. The excerpt names these defenses but provides no design details about token issuance, binding scope, verification procedure, or whether tested attacks actually targeted them.
the trust-boundary part left open in a defense that was measured, so a check at a boundary is not a specified boundary
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A Causal Model for Locating and Unlocking Sandbagging in Model Organisms
- The Troy Moment of AI: Why Some Will Cheat and Some Will Follow?
- Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study
- Memory in the Age of AI Agents: A Survey — Forms, Functions and Dynamics
- Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and Opportunities
- Mining Hidden Thoughts from Texts: Evaluating Continual Pretraining with Synthetic Data for LLM Reasoning
- OpenClaw-RL: Train Any Agent Simply by Talking
- The Addictive Intimacy of AI: Understanding User Disengagement from AI Companions and Why Some Relationships with AI Become Difficult to Leave
Original note title
which parts of the SoK's five-part defense contract do the vault's multi-agent defense notes leave unspecified — a keyword check finds no mention of recovery in five of six