SYNTHESIS NOTE
Topics›Flaws›this note

What cost does making decoys convincing impose on legitimate users?

When honeytokens are designed to look indistinguishable from genuine objects, a total-variation bound constrains how often trusted agents can use the real thing without triggering false alarms. Understanding this tradeoff is key to evaluating whether deceptive security measures protect without harming the people they're meant to spare.

Synthesis note · 2026-09-23 · sourced from Flaws

The abstract's second result reads: "When decoys look like genuine objects, a total-variation bound limits how far legitimate use can be separated from benign activation." The excerpt defines neither "legitimate use" nor "benign activation" and gives no form for the bound.

My reading of the terms, which is not the paper's wording: legitimate use is a trusted agent using a genuine object, and benign activation is a trusted, non-malicious agent setting off a decoy. Total-variation distance measures how far apart two distributions are, and it caps the difference in the probability of any event under them. If that is the bound, then the closer a decoy is to a genuine object, the smaller the gap between how often trusted agents can use the real thing and how often they trip the trap. Making decoys convincing is what makes them worth planting, and the cost falls on the people the trap is meant to spare.

This is the other half of a squeeze. Can honeytokens fool attackers who know the trusted policy? says a rule that separates decoys from genuine objects gets copied. This note says the alternative, removing the difference so no rule can separate them, limits what the trusted side can do with genuine objects. Read together, the design space looks narrow, but that framing is mine. The excerpt states the two results separately and does not present them as a trade-off.

The strongest objection is that a bound is only as useful as it is tight. The excerpt does not say whether it is tight, or how much benign activation a real deployment would tolerate before the alarm stops meaning anything. That tolerance is a false-alert budget, on my reading of a benign activation as a false alert, and Does added monitoring improve protection at acceptable cost? holds such a budget fixed for a different monitor; neither excerpt gives a figure.

What the excerpt does not give. The bound's form and constants, what counts as "looking like" a genuine object, and any measured rate of benign activation.

Inquiring lines that read this note 15

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How can honeytokens stay effective against compromised insider threats? How can we verify agent claims against their actual capabilities and actions? Do planted honeypot tests reliably measure reward hacking? Do current AI defenses adequately protect against semantic manipulation attacks?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 106 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

making decoys look like genuine objects has a price paid by trusted agents — a total-variation bound limits how far legitimate use can be separated from benign activation