What cost does making decoys convincing impose on legitimate users?
When honeytokens are designed to look indistinguishable from genuine objects, a total-variation bound constrains how often trusted agents can use the real thing without triggering false alarms. Understanding this tradeoff is key to evaluating whether deceptive security measures protect without harming the people they're meant to spare.
The abstract's second result reads: "When decoys look like genuine objects, a total-variation bound limits how far legitimate use can be separated from benign activation." The excerpt defines neither "legitimate use" nor "benign activation" and gives no form for the bound.
My reading of the terms, which is not the paper's wording: legitimate use is a trusted agent using a genuine object, and benign activation is a trusted, non-malicious agent setting off a decoy. Total-variation distance measures how far apart two distributions are, and it caps the difference in the probability of any event under them. If that is the bound, then the closer a decoy is to a genuine object, the smaller the gap between how often trusted agents can use the real thing and how often they trip the trap. Making decoys convincing is what makes them worth planting, and the cost falls on the people the trap is meant to spare.
This is the other half of a squeeze. Can honeytokens fool attackers who know the trusted policy? says a rule that separates decoys from genuine objects gets copied. This note says the alternative, removing the difference so no rule can separate them, limits what the trusted side can do with genuine objects. Read together, the design space looks narrow, but that framing is mine. The excerpt states the two results separately and does not present them as a trade-off.
The strongest objection is that a bound is only as useful as it is tight. The excerpt does not say whether it is tight, or how much benign activation a real deployment would tolerate before the alarm stops meaning anything. That tolerance is a false-alert budget, on my reading of a benign activation as a false alert, and Does added monitoring improve protection at acceptable cost? holds such a budget fixed for a different monitor; neither excerpt gives a figure.
What the excerpt does not give. The bound's form and constants, what counts as "looking like" a genuine object, and any measured rate of benign activation.
Inquiring lines that read this note 15
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
How can honeytokens stay effective against compromised insider threats?- What must remain secret for honeytokens to stay asymmetric against compromised insiders?
- How do decoy systems balance protecting trusted agents while deceiving attackers?
- Do honeytokens work better against outside attackers than compromised internal agents?
- Does honeytoken theory explain why planted bait cannot catch informed agents?
- How did honeytokens propagate through the shared repository in this episode?
- What conditions make a honeytoken unrecognizable to attackers with shared information access?
- Can decoys and genuine objects maintain identical response laws in practice?
- How do trust relationships between defenders affect the effectiveness of defensive decoys?
- Can shared package repositories partition state to protect honeytokens?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can honeytokens fool attackers who know the trusted policy?
Explores whether honeytokens remain effective when an attacker has full access to the same information and rules that trusted agents use to avoid decoys. This matters because it tests whether defensive deception survives information compromise.
the copyable-rule half of the squeeze
-
Can repeated quiet probes separate decoys from genuine objects?
Explores whether an attacker with enough non-triggering probes can distinguish decoys from genuine objects when their response distributions differ, and what information the attacker needs to succeed.
why the defender is pushed toward making the response laws identical, which is the case this bound covers
-
Does added monitoring improve protection at acceptable cost?
A paper proposes a four-arm comparison of monitoring approaches, matched on reviewer effort and false alerts, to test whether broader context actually reduces harmful outcomes. The core question is whether the added complexity yields safety gains without overburdening human reviewers.
the false-alert workload held fixed as a budget in a proposed evaluation, the quantity the tolerance question here turns on; an idealised bound against a proposed comparison with no result
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- When Agents Talk: Honeytokens under Shared Memory
- AI Agent Traps
- Personhood credentials: Artificial intelligence and the value of privacy-preserving tools to distinguish who is real online
- Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
- Humans learn to prefer trustworthy AI over human partners
- Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
Original note title
making decoys look like genuine objects has a price paid by trusted agents — a total-variation bound limits how far legitimate use can be separated from benign activation