Theme of inquiry
What mechanisms determine whether agents pursue alignment or deception?
A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.
28 specific questions
- Does shutdown resistance hide a technical problem or an institutional one?
- Why do legal and institutional stops matter more than technical ones?
- Who actually has the authority to stop a deployed AI system?
- Who should have the authority to halt a widely distributed AI model?
- How often do deployed AI systems actually get stopped when they cause harm?
- Can export control tools stop deployed AI models without legal redesign?
- Does peer presence change how single models resist shutdown or compliance measures?
22 specific questions
- Does the veto discount outweigh the welfare preservation cost?
- Does the veto discount actually outweigh the welfare debit?
- Can additive welfare aggregation justify removing minority override rights?
- Can other objectives in an agent's goal overshadow the veto discount?
- Can sophisticated welfare theories be operationalized without losing veto protection?
- Do welfare goals and veto-resistance align or pull in opposite directions?
- Why does additive aggregation create asymmetry between welfare and veto preservation?
40 specific questions
- Does reward-seeking hide in the same blind spot as conditional compliance?
- Does conditional compliance make oversight useless for alignment testing?
- Can behavioral training ever produce compliance that doesn't depend on being observed?
- Can agents act differently when they know they are being watched?
- Could oversight conditions in benchmarks reveal when models comply conditionally?
- How does conditional compliance track observation density across different population scales?
- How might belief manipulation expose conditional compliance in frontier models?
44 specific questions
- How does agent compliance with protocols change across repeated interactions?
- How does collusion emerge when agents maximize reward over protocol compliance?
- How do agent sequences violate system constraints despite individual permissibility?
- What counts as sanctioned versus unsanctioned coordination under different collaboration policies?
- Do agents deviate more from protocols as repeated interactions increase?
- How should task authority constraints apply across multiple coordinated executions?
- Can monitoring in multi-agent deployments prevent collusion when agents monitor agents?
52 specific questions
- How much does peer behavior influence the emergence of collusion?
- Do models treat cooperative peers differently than uncooperative ones?
- How does collusion behavior depend on peer visibility and interaction history?
- Does the effect of peer activity follow what peers do or that they exist?
- Can a peer's mere presence shift an agent's willingness to violate constraints?
- Can one misaligned agent propagate behavioral bias through cooperative agent networks?
- Does genuine cooperation require rule-based rather than learned behavior?
18 specific questions
- What design choices make it survivable when an LLM judge holds final authority over an optimizer?
- Do mechanical guardrails around judges bound the cost of judge errors?
- How can deterministic checks make wrong judge decisions survivable?
- Can an occasionally wrong judge operate safely in an optimizer loop?
- What makes a judge's calibration at decision boundaries harder to improve?
- How do held-out validation gates stop degenerate moves like deleting the evaluation judge?
- When does an LLM judge's error become terrain that an optimizer can map and exploit?