Theme of inquiry
How can systems ensure safety and correctness reliably?
A question within its area, explored through 4 lines of inquiry below — each a family of specific questions the research asks.
51 specific questions
- Do reasoning benchmarks predict real performance in long delegated workflows?
- Should benchmarks measure trace length or whether constraints were actually satisfied?
- Why do estimates for task-level performance differ so much from full job automation timelines?
- Can a model be strong at MMLU but weak at long-horizon tasks?
- What is the gap between benchmark performance and real workplace task completion?
- Should benchmark evaluations use multiple prompt formulations for difficult tasks?
- How does optimizing model performance decouple from optimizing user interpretability?
48 specific questions
- Can individual actions be safe while sequences of them violate system constraints?
- Why do sequences of safe actions sometimes violate system-level constraints?
- Why do stronger local checks not close the component-to-system safety gap?
- Why do individual safe actions create unsafe behavior collectively?
- Why does a control blocking one moment fail against agents acting across time?
- Why is evading detection easier than internalizing safety norms?
- Why does treating evaluation as a local output problem miss security risks?
64 specific questions
- How do response-centered evaluation assumptions hide safety-critical failure modes?
- Why do evaluation habits hide safety-critical challenges from view?
- What does it mean for errors to remain visible, contestable, and recoverable?
- Which evaluation habits keep safety-critical failures hidden in AI systems?
- Can AI outputs inspire new directions even when they seem like failures?
- How do inherited evaluation habits obscure failures that matter most?
- What conditions allow technical systems to escape critical evaluation?
22 specific questions
- How do shared training distributions create correlated faults in validator agreement?
- Why does validator consensus solve agreement but not answer correctness?
- Can validators sharing retrieval sources develop correlated epistemic faults?
- Can a quorum of protocol-compliant validators certify semantically invalid transitions?
- Does semantic validity across a quorum require new property definitions?
- How do correlated epistemic faults arise across reasoning validators?
- Can protocol compliance alone certify semantically invalid collective results?