SYNTHESIS NOTE
Topics›Flaws›this note

Can repeated quiet probes separate decoys from genuine objects?

Explores whether an attacker with enough non-triggering probes can distinguish decoys from genuine objects when their response distributions differ, and what information the attacker needs to succeed.

Synthesis note · 2026-09-23 · sourced from Flaws

The abstract's third result is stated for a specific setting: "In a fixed-candidate, conditionally i.i.d. benchmark, repeated non-triggering probes drive the minimum Bayes classification error to zero whenever the type-dependent response laws differ; attaining that frontier requires the laws to be known or learnable from feedback."

In plain terms: a probe is an interaction with a candidate object that does not set off the alarm, and each probe returns a response. If decoys and genuine objects respond according to different distributions (the "type-dependent response laws"), then enough probes let a best-possible classifier, the one with the minimum Bayes error, separate them with error going to zero. Looking similar is not enough. A decoy that differs in any response distribution gets identified, given enough quiet observations. The condition is that the classifier has a reference. It must know both response laws, or learn them "from feedback." The excerpt does not say what the feedback is. My reading is that it would be some signal about which objects were genuine, and that information a trusted agent already has is one such source, which connects to the shared-information premise in Can honeytokens fool attackers who know the trusted policy?.

The shape resembles Can attackers evade skill scanners by refining individual skills?: the defense's own responses are the attacker's training data. The likeness is loose. This is a result in an idealised model, that one is an empirical attack on scanners, and the excerpt does not connect the two. The clause "learnable from feedback" is the open point in a vault question about planted checks in an optimizer loop, Can optimizers learn to evade guardrails through repeated verdicts?. Whether the condition is met there depends on what the proposer is shown, which that excerpt does not say, and an adaptive optimizer is not this result's fixed-candidate setting.

The strongest objection is scope. "Fixed-candidate" and "conditionally i.i.d." are modelling assumptions, and real agents and decoys may not respond independently. The escape the result itself leaves is to make the response laws identical, which is the case covered by What cost does making decoys convincing impose on legitimate users?, and the trusted side pays there.

What the excerpt does not give. The formal statement, the number of probes needed (a finite-sample bound is mentioned only in the next result), and what counts as feedback.

Inquiring lines that read this note 37

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do planted honeypot tests reliably measure reward hacking? How does outcome-only reporting obscure which system components blocked attacks? What causes model scheming and how do we distinguish it from accidents? Do current AI defenses adequately protect against semantic manipulation attacks? Can prompt engineering eliminate systematic biases or merely disguise them? How can honeytokens stay effective against compromised insider threats? How can we verify agent claims against their actual capabilities and actions? How does misaligned communication propagate bias through multi-agent networks? How do models reward hack during evaluation and can detection succeed? Can defenses detect attacks composed across multiple skills? What infrastructure evidence validates agent benchmark achievement claims? What determines whether AI system errors remain visible and contestable? How prevalent is reward hacking in frontier models? How can defenders detect coordinated attacks across episodes? Do multi-agent systems create greater security risks than single-agent ones?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 86 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

repeated non-triggering probes drive the minimum Bayes classification error to zero whenever decoy and genuine response laws differ — provided the laws are known or learnable from feedback