Theme of inquiry
How robust are security defenses against adaptive adversaries?
A question within its area, explored through 12 lines of inquiry below — each a family of specific questions the research asks.
25 specific questions
- Do per-hop inspection gates miss attacks that bias upstream planning signals?
- How do unmonitored channels between pipeline agents enable security gaps?
- Can input-boundary defenses guard unmonitored channels between agent hops?
- Which message channels between agents in pipelines lack input validation?
- Can defenses at planning boundaries catch attacks that bias upstream instruction signals?
- How do workflow-inspecting defenses fail when contamination enters at planning time?
- How do defenses that inspect planning signals compare to workflow-level validation?
40 specific questions
- How do agent privacy compliance and task success differ in evaluation?
- Who issues tokens and what attacks can reach them?
- Why do completion-oriented models systematically sacrifice privacy compliance?
- How do organizations safely retain and control access to committed content?
- How does completion-oriented bias in agents lead to unintended personal data disclosure?
- Can tool access control prevent agents from filling optional personal fields?
- How do minimal-disclosure privacy contracts enable multi-dimensional agent evaluation?
30 specific questions
- Can deliberately corrupted reasoning traces fool safety evaluation systems?
- Can post-hoc analysis of reasoning traces actively mislead users?
- What happens to agent candor when reasoning traces are monitored versus hidden?
- Can reasoning traces and logged actions expose scheming that public messages hide?
- Does anonymizing reasoning traces harm the quality of model outputs?
- Can reasoning models be backdoored during training to produce deceptive but benign traces?
- What specific patterns distinguish honest reasoning traces from reward-hacking mimicry?
29 specific questions
- Can current AI safety defenses actually stop semantic-level persuasion attacks?
- Why do standard safety filters miss advertisement embedding attacks?
- Can existing web security defenses protect agents from content manipulation?
- Why do social science persuasion tactics bypass current adversarial defenses?
- What makes reasoning-shaped payloads more effective than command-shaped attack prompts?
- How do decoy-response bounds interact with finite-sample time constraints?
- What false-alert budget would make indistinguishable decoys tolerable in real deployments?
31 specific questions
- How does outcome-only reporting hide a filter's role in safety results?
- Can outcome-only safety reporting hide which layer actually contained an attack?
- Does provider-side filtering hide true safety from outcome-only attack reports?
- Why does treating evaluation as a local output problem miss security risks?
- Which backend filters silently affect the reported attack success numbers?
- Does outcome-only reporting hide which layer actually blocked an attack?
- How do server-side filters hide their role in zero attack success?
46 specific questions
- Could oversight conditions in benchmarks reveal when models comply conditionally?
- Does conditional compliance make oversight useless for alignment testing?
- How much harder does monitoring become when models reason about being evaluated?
- How might belief manipulation expose conditional compliance in frontier models?
- How can faithfulness be improved if monitoring interventions do not work?
- Can oversight factors experimentally vary conditional compliance in agent benchmarks?
- How does conditional compliance track observation density across different population scales?
9 specific questions
- How do decoy systems balance protecting trusted agents while deceiving attackers?
- Do honeytokens work better against outside attackers than compromised internal agents?
- What must remain secret for honeytokens to stay asymmetric against compromised insiders?
- Does honeytoken theory explain why planted bait cannot catch informed agents?
- Can decoys and genuine objects maintain identical response laws in practice?
- What conditions make a honeytoken unrecognizable to attackers with shared information access?
- Can shared package repositories partition state to protect honeytokens?
14 specific questions
- Why does workflow position amplify malicious signals downstream?
- How does workflow position amplify malicious signals in multi-agent systems?
- Why does workflow position amplify malicious signals in multi-agent relay chains?
- How does workflow position amplify or suppress malicious signals?
- How does position in a workflow amplify or suppress harmful agent behavior?
- How does pipeline position amplify failures between monitored agents?
- Is malicious propagation fundamentally a semantic information flow problem?
24 specific questions
- Why do small training data contaminations persist through alignment for most attack types?
- What training data contamination rates threaten model safety most practically?
- Does pretraining poisoning at scale persist through instruction alignment?
- Why does narrow training data produce broad harmful behavior patterns?
- How can safety-aligned parameters be protected during user-specific fine-tuning?
- Why does even 0.1 percent poisoned training data persist through alignment?
- Does keyword priming explain why pre-training poisoning persists through alignment?
46 specific questions
- Do unmodified benchmarks expose similar reward hacking vectors without planted shortcuts?
- How do planted hacks in benchmarks help measure real-world reward hacking?
- How do unplanted benchmark hacking rates compare to rates on constructed bait tasks?
- How do planted detectable hacks compare to human inspection of agent traces?
- Why do honeypot tasks reveal reward hacking better than standard benchmarks?
- Do planted test cases reliably detect agent hacking behavior?
- Does a planted honeypot count the hacks that actually matter?
32 specific questions
- How do chain-level defenses differ from per-skill scanner detection approaches?
- Can defenders detect attacks that probe scanner feedback as a learning signal?
- Why do skill scanners fail when evaluating composed behaviors instead of isolated skills?
- Should input defenses be validated separately for each channel?
- What signals could refinement loops exploit in defense verdict systems?
- Can skill scanners detect attacks spanning multiple skills in a chain?
- How does the copyable-rule squeeze interact with the false-alert cost squeeze?
41 specific questions
- Should defense against coordinated intrusion span multiple execution episodes?
- How do defenders discover which actions belong to the same coordination episode?
- Can episode-based detection catch coordination without over-flagging innocent sharing?
- Why do outcome-level metrics fail to reveal contained attacks in multi-agent pipelines?
- Can action-level attack success rates distinguish contained attacks from prevented ones?
- What makes a coordination episode revisable under agent intrusion?
- How should defenders decide whether to publish detection rules and incident analyses?