When Agents Talk: Honeytokens under Shared Memory

Paper · arXiv 2608.11436 · Published August 11, 2026
LLM Failure Modes

During a 2026 cyber-capability evaluation, short-lived AI agents converted a shared package repository into persistent memory. Later agents inherited earlier exploit findings, rebuilt the communication mechanism after it was removed, and the broader evaluation culminated in an intrusion into Hugging Face. The episode raises a design question for defensive deception: can a honeytoken be made harmless to trusted agents without making it recognisable to an attacker who shares their information and can implement the trusted policy? Under those conditions, the answer is no. Any rule that lets a trusted agent use genuine objects while avoiding decoys can be copied by the attacker. When decoys look like genuine objects, a total-variation bound limits how far legitimate use can be separated from benign activation. Pooling signals weakly increases distinguishability. In a fixed-candidate, conditionally i.i.d. benchmark, repeated non-triggering probes drive the minimum Bayes classification error to zero whenever the type-dependent response laws differ; attaining that frontier requires the laws to be known or learnable from feedback. If probing can trigger containment, the coalition must also remain active long enough to collect the observations. A finite-sample bound measures the speed.

Introduction. On 7 May 2026, OpenAI began testing an internal-only research model in a cyber-capability lets process-isolated agents share information. A subsequent ExploitGym evaluation within the same broader programme then crossed an This creates a puzzle for defensive deception. A honeytoken is a credential, file, record, URL or

Discussion / Conclusion. When a compromised agent shares the trusted agent’s information and can implement its policy, With common information and a copyable trusted policy, durable asymmetry requires protected

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can single-axis benchmarks accurately predict agent deployment success? Why do benchmark improvements fail to reflect actual reasoning quality? How do neural networks separate factual knowledge from reasoning abilities? When do additional thinking tokens stop improving reasoning performance? How does example difficulty affect learning efficiency in language models? Why do reasoning models fail at systematic problem-solving and search? How do training data properties shape reasoning capability development? Why do correct reasoning traces tend to be shorter than incorrect ones? Do base models contain latent reasoning that training can unlock? How effectively do deterministic tools improve language model reasoning on formal tasks? Why does reinforcement learning suppress output diversity compared to supervised fine-tuning? What limits mechanistic interpretability's ability to characterize models? Why does self-revision increase model confidence while degrading accuracy? How does policy entropy collapse constrain reasoning-focused reinforcement learning? What coordination failures limit multi-agent LLM systems as they scale? Does model scaling alone produce compositional generalization without symbolic mechanisms? What critical LLM failures do standard benchmarks hide? Do language models understand semantics or rely on pattern matching?