Theme of inquiry

How robust are security defenses against adaptive adversaries?

A question within its area, explored through 12 lines of inquiry below — each a family of specific questions the research asks.


Where do unmonitored channels leave multi-agent planning vulnerable to attack?

25 specific questions

See all 25 questions in this line of inquiry
How do agents balance task completion with privacy compliance and security?

40 specific questions

See all 40 questions in this line of inquiry
How reliable are reasoning traces as evidence of agent honesty?

30 specific questions

See all 30 questions in this line of inquiry
Do current AI defenses adequately protect against semantic manipulation attacks?

29 specific questions

See all 29 questions in this line of inquiry
How does outcome-only reporting obscure which system components blocked attacks?

31 specific questions

See all 31 questions in this line of inquiry
How can evaluations detect conditional compliance in monitored AI systems?

46 specific questions

See all 46 questions in this line of inquiry
How can honeytokens stay effective against compromised insider threats?

9 specific questions

See all 9 questions in this line of inquiry
How does position in multi-agent workflows amplify or attenuate harmful signals?

14 specific questions

See all 14 questions in this line of inquiry
How does training data contamination persist through safety alignment mechanisms?

24 specific questions

See all 24 questions in this line of inquiry
Do planted honeypot tests reliably measure reward hacking?

46 specific questions

See all 46 questions in this line of inquiry
Can defenses detect attacks composed across multiple skills?

32 specific questions

See all 32 questions in this line of inquiry
How can defenders detect coordinated attacks across episodes?

41 specific questions

See all 41 questions in this line of inquiry