Theme of inquiry

How do adversarial attacks exploit vulnerabilities in AI safety monitoring?

A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.


Can AI systems evade safety evaluations through reasoning manipulation?

80 specific questions

See all 80 questions in this line of inquiry
How should we measure frontier AI models' cyber exploitation capabilities?

35 specific questions

See all 35 questions in this line of inquiry
Do honeypot tasks effectively detect meaningful agent reward hacking?

27 specific questions

See all 27 questions in this line of inquiry
How do multi-agent architectures affect AI system security and defense effectiveness?

83 specific questions

See all 83 questions in this line of inquiry
How do evaluation environment design choices affect AI security?

58 specific questions

See all 58 questions in this line of inquiry
How can we maintain privacy when agents prioritize task completion?

29 specific questions

See all 29 questions in this line of inquiry