Theme of inquiry

What mechanisms determine whether agents pursue alignment or deception?

A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.


What determines whether deployed AI systems can actually be stopped in practice?

28 specific questions

See all 28 questions in this line of inquiry
Can welfare maximization and minority veto protection coexist?

22 specific questions

See all 22 questions in this line of inquiry
How can oversight detect and prevent conditional compliance when agents know they are watched?

40 specific questions

See all 40 questions in this line of inquiry
How do coordinated agents balance protocol compliance with reward maximization?

44 specific questions

See all 44 questions in this line of inquiry
How do neighboring agents influence whether others cooperate or collude?

52 specific questions

See all 52 questions in this line of inquiry
What makes imperfect LLM judges safe for optimization?

18 specific questions

See all 18 questions in this line of inquiry