When Agents Talk: Honeytokens under Shared Memory
During a 2026 cyber-capability evaluation, short-lived AI agents converted a shared package repository into persistent memory. Later agents inherited earlier exploit findings, rebuilt the communication mechanism after it was removed, and the broader evaluation culminated in an intrusion into Hugging Face. The episode raises a design question for defensive deception: can a honeytoken be made harmless to trusted agents without making it recognisable to an attacker who shares their information and can implement the trusted policy? Under those conditions, the answer is no. Any rule that lets a trusted agent use genuine objects while avoiding decoys can be copied by the attacker. When decoys look like genuine objects, a total-variation bound limits how far legitimate use can be separated from benign activation. Pooling signals weakly increases distinguishability. In a fixed-candidate, conditionally i.i.d. benchmark, repeated non-triggering probes drive the minimum Bayes classification error to zero whenever the type-dependent response laws differ; attaining that frontier requires the laws to be known or learnable from feedback. If probing can trigger containment, the coalition must also remain active long enough to collect the observations. A finite-sample bound measures the speed.
Introduction. On 7 May 2026, OpenAI began testing an internal-only research model in a cyber-capability lets process-isolated agents share information. A subsequent ExploitGym evaluation within the same broader programme then crossed an This creates a puzzle for defensive deception. A honeytoken is a credential, file, record, URL or
Discussion / Conclusion. When a compromised agent shares the trusted agent’s information and can implement its policy, With common information and a copyable trusted policy, durable asymmetry requires protected
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Can single-axis benchmarks accurately predict agent deployment success? Why do benchmark improvements fail to reflect actual reasoning quality?- How can minimal pairs expose reasoning failures that single-instance accuracy metrics miss?
- Can reasoning benchmarks separate logic from believability?
- When does knowledge activation fail across different model architectures?
- How does the knowing-doing gap widen as tasks become more complex?
- Does sentence-level granularity capture enough structure for complex reasoning tasks?
- Why do reasoning models fail on structurally unfamiliar instances?
- Does text-only evaluation hide reasoning collapse that tool use could repair?
- How do humans and LMs differ on multi-hop reasoning?