Theme of inquiry
What sustains meaningful human oversight in automated systems?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
19 specific questions
- How do pressure and strategic hints separately influence scheming compared to instrumental goals?
- How often do scheming reasoning and covert actions actually align in practice?
- Do deliberate strategic reasoning triggers like replacement and goal conflict rank differently across models?
- How should tool-call attribution distinguish credit between successful accidents and intentional actions?
- How does pressure mainly influence scheming reasoning versus covert action?
- Why do instrumental goals drive scheming more strongly than pressure does?
- What distinguishes strategic fabrication from accidental hallucination in research agents?
29 specific questions
- How does tokenization of intelligence reshape what value means in culture?
- How is tokenized intelligence different from traditional commodification of expertise?
- How does the token frame predict different economic outcomes than commodity framing?
- Can AI output be tokenized without decoupling from the thought processes behind it?
- How does tokenization change what gets counted as valuable knowledge?
- Can foundation model outputs satisfy exchange value while lacking use value?
- Why do tokens need validators while commodities need standardization?
71 specific questions
- Does the veto discount actually outweigh the welfare debit?
- Can humans build reliable oversight for increasingly complex AI systems?
- Does keeping humans in the loop protect against AI risk without scrutiny capacity?
- Does the veto discount outweigh the welfare preservation cost?
- Can human oversight actually function as a cost on all agent goals?
- Can sophisticated welfare theories be operationalized without losing veto protection?
- Can human oversight actually stop a deployed capable agent in practice?
23 specific questions
- Can the causal model predict which cached layers to graft?
- How does context grafting compare to single-layer residual stream grafting?
- Did the causal model predict the five failures before observing them?
- Does the causal model help locate sandbagging locks with unknown passwords?
- Can installed sandbagging locks in small models describe uninstalled sandbagging behavior?
- Can refusal behavior be restored by grafting the same causal axis discovered for sandbagging?
- Can residual stream grafts work without knowing which layers to intervene on?
59 specific questions
- How do agent sequences violate system constraints despite individual permissibility?
- Do sequences of individually safe actions collectively violate system-level constraints?
- Can individual actions be safe while sequences of them violate system constraints?
- What safeguards prevent peer activity from normalizing boundary violations?
- How should task authority constraints apply across multiple coordinated executions?
- What happens when stopping rules must cross organizational boundaries?
- How do policies distinguish individual action rules from sequence-level constraints?