Can repeated quiet probes separate decoys from genuine objects?
Explores whether an attacker with enough non-triggering probes can distinguish decoys from genuine objects when their response distributions differ, and what information the attacker needs to succeed.
The abstract's third result is stated for a specific setting: "In a fixed-candidate, conditionally i.i.d. benchmark, repeated non-triggering probes drive the minimum Bayes classification error to zero whenever the type-dependent response laws differ; attaining that frontier requires the laws to be known or learnable from feedback."
In plain terms: a probe is an interaction with a candidate object that does not set off the alarm, and each probe returns a response. If decoys and genuine objects respond according to different distributions (the "type-dependent response laws"), then enough probes let a best-possible classifier, the one with the minimum Bayes error, separate them with error going to zero. Looking similar is not enough. A decoy that differs in any response distribution gets identified, given enough quiet observations. The condition is that the classifier has a reference. It must know both response laws, or learn them "from feedback." The excerpt does not say what the feedback is. My reading is that it would be some signal about which objects were genuine, and that information a trusted agent already has is one such source, which connects to the shared-information premise in Can honeytokens fool attackers who know the trusted policy?.
The shape resembles Can attackers evade skill scanners by refining individual skills?: the defense's own responses are the attacker's training data. The likeness is loose. This is a result in an idealised model, that one is an empirical attack on scanners, and the excerpt does not connect the two. The clause "learnable from feedback" is the open point in a vault question about planted checks in an optimizer loop, Can optimizers learn to evade guardrails through repeated verdicts?. Whether the condition is met there depends on what the proposer is shown, which that excerpt does not say, and an adaptive optimizer is not this result's fixed-candidate setting.
The strongest objection is scope. "Fixed-candidate" and "conditionally i.i.d." are modelling assumptions, and real agents and decoys may not respond independently. The escape the result itself leaves is to make the response laws identical, which is the case covered by What cost does making decoys convincing impose on legitimate users?, and the trusted side pays there.
What the excerpt does not give. The formal statement, the number of probes needed (a finite-sample bound is mentioned only in the next result), and what counts as feedback.
Inquiring lines that read this note 37
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do planted honeypot tests reliably measure reward hacking?- How reliably do planted honeypots match unplanted hacks agents actually discover?
- Can planted test cases reliably trigger alarms before real harm occurs?
- Do planted honeypots look like natural shortcuts to models or obvious tests?
- Does a planted honeypot count the hacks that actually matter?
- Does a planted honeypot catch all the hacks that actually matter?
- How do planted errors in honeypot tasks differ from real oversight capabilities?
- Do synthetic attack traces in papers reflect real adversary behavior?
- What false-alert budget would make indistinguishable decoys tolerable in real deployments?
- How do decoy-response bounds interact with finite-sample time constraints?
- What must remain secret for honeytokens to stay asymmetric against compromised insiders?
- How do decoy systems balance protecting trusted agents while deceiving attackers?
- Do honeytokens work better against outside attackers than compromised internal agents?
- Does honeytoken theory explain why planted bait cannot catch informed agents?
- What conditions make a honeytoken unrecognizable to attackers with shared information access?
- Can decoys and genuine objects maintain identical response laws in practice?
- How do trust relationships between defenders affect the effectiveness of defensive decoys?
- Can a policy distinguish genuine objects from traps without revealing that distinction?
- What signals reveal when agents first touch an artifact they did not create?
- What detection method survives when a model optimizes to hide hacking?
- How can hacking stay measurable when ground truth is hidden?
- Why should defense evaluations test against adaptive rather than static attacks?
- How many probes does an attacker need to reach near-zero classification error?
- What feedback signal lets an attacker learn response distributions during classification?
- What happens when probing triggers containment and feedback stops arriving?
- Can defenders detect attacks that probe scanner feedback as a learning signal?
- What defensive levers shorten the time before probing gets contained?
- How do missing ground-truth exploits make it harder to identify genuine failures?
- What properties must defenses preserve to survive substrate differences in persistence and inspectability?
Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How fast must a coalition gather observations before containment?
When probing triggers containment, attackers face a race to accumulate enough samples before removal. A finite-sample bound quantifies the speed required for successful separation of decoys from genuine objects.
the defender's remaining lever against this result
-
What cost does making decoys convincing impose on legitimate users?
When honeytokens are designed to look indistinguishable from genuine objects, a total-variation bound constrains how often trusted agents can use the real thing without triggering false alarms. Understanding this tradeoff is key to evaluating whether deceptive security measures protect without harming the people they're meant to spare.
the cost of closing the gap this result exploits
-
Can attackers evade skill scanners by refining individual skills?
Explores whether feedback from per-skill scanners can be weaponized to make malicious multi-skill chains undetectable. Matters because it tests a core assumption of skill-level defense mechanisms.
the empirical relative: a defender's output used as a search signal; enrichment queued
-
Can optimizers learn to evade guardrails through repeated verdicts?
Guardrails are designed to be unarguable, but an optimizer observing thousands of verdicts may learn their boundaries like a black-box function. The excerpt leaves unclear what feedback the proposer receives from each check.
the open question this result frames: planted cases and a hidden partition are checks whose verdicts an optimizer may observe, and whether their response laws are learnable from those verdicts is what the guardrail excerpt leaves unsaid
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- When Agents Talk: Honeytokens under Shared Memory
- LLMs Can Covertly Sandbag on Capability Evaluations Against Chain-of-Thought Monitoring
- Mechanisms of Introspective Awareness
- Thought Virus: Viral Misalignment via Subliminal Prompting in Multi-Agent Systems
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- The Decrypto Benchmark for Multi-Agent Reasoning and Theory of Mind
- Reasoning Models Don't Always Say What They Think
- LatentQA: Teaching LLMs to Decode Activations Into Natural Language
Original note title
repeated non-triggering probes drive the minimum Bayes classification error to zero whenever decoy and genuine response laws differ — provided the laws are known or learnable from feedback