Who decides which agent communications get anchored?
The paper commits to anchoring 'selected' communications but never specifies who makes that selection, by what criteria, or how missed selections would be detected. This matters because the selector controls what evidence can ever exist.
Two words carry the design. The abstract anchors commitments for "selected agent communications," and the conclusion speaks of "critical process traces." The GRC list adds "risk-based evidence selection" as a practice. Anchoring everything is presumably not the goal, since cost and privacy would both object. That is my reading, and the excerpt does not say why the selection is there.
If selection is by design, then selection is itself a control. Whatever is not selected has no anchor and can never later be shown unmodified. So the open questions are about the selector. Who chooses: a compliance function, the platform, the agents themselves? By what risk criteria? Fixed in advance or adaptive? And how is a miss detected, given that a missed trace leaves no anchor to notice its absence? The evidence model's "authorized anchoring" (What can a blockchain anchor actually prove about records?) asks who may anchor, and selection asks who decides what is worth anchoring. The excerpt does not say whether these are one control or two.
A risk-based selector predicts what will matter. The incident the paper cites as motivation reportedly involved traffic through unauthorized channels (Can a black box see communication through unauthorized channels?), which is the kind of traffic an expectation-based selector is least likely to list. This is a worry and not a finding, since the excerpt does not describe the selection method at all.
Two neighbours in the vault bear on the selector, both from other settings and neither on this paper. What to keep depends on which actions belong together, and Can defenders discover agent episodes without knowing membership in advance? asks how to find that grouping before anyone hands it over. Can a finite lifecycle model detect reward hacking across benchmarks? writes down what to record as a finite typed lifecycle of reward-relevant events. My reading is that such a rule can be enumerated because a benchmark's reward path is closed, and the excerpt does not say who writes a binding. Whether an open agent workflow has a comparable finite model is addressed by neither excerpt.
The vault also holds one device for making a control's silent miss visible: the fourth move in Can deterministic checks protect LLM judges from failure? plants a known case and treats its failure to register as the alarm. Applied to a selector it would mean planting a known communication and checking that it gets anchored. That transfer is my suggestion and the black-box excerpt proposes nothing like it. It would test the selector only on what someone thought to plant, which is the coverage limit the vault records for planted cases, and the traffic an expectation-based selector is least likely to list is the traffic least likely to be planted.
A usable test for anyone writing about the layer: before saying it "would have shown" something, ask whether that something would have been selected.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How do agents balance task completion with privacy compliance and security? What infrastructure evidence validates agent benchmark achievement claims? How does misaligned communication propagate bias through multi-agent networks? How can we verify agent claims against their actual capabilities and actions?Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
What can a blockchain anchor actually prove about records?
Blockchain anchors provide tamper evidence, but the note explores what properties they cannot guarantee—like whether events occurred in the right order, were captured accurately, or were authorized to be anchored in the first place.
"authorized anchoring" is the nearest named property to a selection authority
-
Can a black box see communication through unauthorized channels?
The black box architecture records sanctioned agent communications, but the paper doesn't specify where capture occurs or whether it detects traffic outside authorized channels. This matters for evaluating whether the system would have recorded the incident that motivated it.
the case where expectation-based selection would most plausibly miss
-
Does anchored evidence actually enable regulatory compliance or just readiness?
The paper proposes blockchain-anchored evidence for five governance uses under three EU regimes, but leaves unclear whether the evidence layer closes the gap between audit readiness and actual compliance. What architectural controls remain unmapped?
where risk-based selection appears in the GRC list
-
Can defenders discover agent episodes without knowing membership in advance?
The core challenge in defending against coordinated agent intrusions is grouping actions into episodes before any external authority labels them. Current methods lack clear discovery techniques, and the trade-off between detection accuracy and reviewer workload remains unresolved.
a selection problem of the same kind on the defence side: what to keep depends on grouping actions before a grouping is supplied; that note already links here
-
Can a finite lifecycle model detect reward hacking across benchmarks?
Does modeling benchmark runs as typed event lifecycles, checked against task bindings, successfully detect reward-hacking exploits across multiple evaluation tasks? This approach aims to replace task-specific patches with reusable formal detection.
one written-down rule for what to record, in a closed benchmark setting; who writes the bindings is open there as who selects is open here
-
Can deterministic checks protect LLM judges from failure?
Explores whether mechanical, non-contestable verification steps can safeguard LLM-based decision systems. Matters because it tests whether we can make AI judgment survivable even when it goes wrong.
planted cases as a candidate way to see a selector's miss, my transfer and untested; it covers only what was planted
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A Black Box for Agentic Processes: Blockchain-Anchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits
- The Missing Layer of AGI: From Pattern Alchemy to Coordination Physics
- Counter-Swarm Doctrine: Containing Coordinated Agent Intrusions
- The Troy Moment of AI: Why Some Will Cheat and Some Will Follow?
- Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce
- Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
- SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
- Artifacts as Memory Beyond the Agent Boundary
Original note title
how is it decided which agent communications get anchored — the architecture anchors selected traces and the GRC discussion names risk-based evidence selection but the excerpt does not say how selection is made or checked