SYNTHESIS NOTE
Topics›Flaws›this note

Can honeytokens fool attackers who know the trusted policy?

Explores whether honeytokens remain effective when an attacker has full access to the same information and rules that trusted agents use to avoid decoys. This matters because it tests whether defensive deception survives information compromise.

Synthesis note · 2026-09-23 · sourced from Flaws

The abstract poses a design question for defensive deception: "can a honeytoken be made harmless to trusted agents without making it recognisable to an attacker who shares their information and can implement the trusted policy?" Its answer is flat: "Under those conditions, the answer is no." The reason comes in one sentence: "Any rule that lets a trusted agent use genuine objects while avoiding decoys can be copied by the attacker." The introduction begins to define a honeytoken as "a credential, file, record, URL or" and then breaks off.

The reasoning is about where a trap's asymmetry comes from. A honeytoken is dangerous only to whoever touches it, so trusted agents need some way to steer around it or they will trip the alarm in ordinary work. That steering is a rule: use this, avoid that. The abstract's premise is that the attacker has the same information and can run the same rule, so whatever spares the trusted agent also spares the attacker. My reading is that the trusted policy is itself the tell. The two requirements, harmless to the trusted and unrecognisable to the attacker, can only both hold if the rule that separates decoys from genuine objects cannot be reproduced, and the stated conditions take that away.

The claim is conditional and should be quoted that way. It says nothing about an outside attacker who lacks the information or cannot implement the policy. The conclusion frames the threat as "a compromised agent" that "shares the trusted agent's information and can implement its policy," and its next sentence stops at the word "protected" (What must honeytokens protect to stay undetectable?). The strongest objection is that this leaves honeytokens intact against the more common outsider case, so it is a limit on one threat model, not a verdict on the technique.

What the excerpt does not give. The formal statement of the result, how "harmless" and "trusted policy" are defined, the proof, and any honeytoken actually deployed in the episode the abstract describes.

Inquiring lines that read this note 18

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do planted honeypot tests reliably measure reward hacking? How does outcome-only reporting obscure which system components blocked attacks? How can honeytokens stay effective against compromised insider threats? Do multi-agent systems create greater security risks than single-agent ones? How do agents balance task completion with privacy compliance and security? How can defenders detect coordinated attacks across episodes? Can defenses detect attacks composed across multiple skills?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
16 direct connections · 132 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

a honeytoken cannot be harmless to trusted agents and unrecognisable to an attacker who shares their information and can implement the trusted policy — any rule that avoids decoys can be copied