Theme of inquiry

How can effective AI defenses withstand adaptive adversarial attacks?

A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.


What attack surfaces do reasoning traces and chains introduce?

48 specific questions

See all 48 questions in this line of inquiry
Can single-point security defenses protect multi-agent systems from multi-step attacks?

70 specific questions

See all 70 questions in this line of inquiry
Do backend defenses obscure real attack effectiveness in reported metrics?

26 specific questions

See all 26 questions in this line of inquiry
How can we detect and prevent harm propagation through multi-agent delegation workflows?

32 specific questions

See all 32 questions in this line of inquiry
Can causal models help detect and locate hidden sandbagging in AI?

26 specific questions

See all 26 questions in this line of inquiry
How effective are honeytokens and decoys against different security threats?

17 specific questions

See all 17 questions in this line of inquiry
How vulnerable are token issuance and authorization policies to coordinated attacks?

14 specific questions

See all 14 questions in this line of inquiry
How do we enforce security boundaries in evaluation environments?

45 specific questions

See all 45 questions in this line of inquiry