Theme of inquiry
How do architectural choices affect system safety and transparency?
A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.
58 specific questions
- Why does infrastructure-side evidence matter more than agent-reported traces?
- Can infrastructure evidence ground benchmark claims better than terminal scores alone?
- What makes recorded transitions more trustworthy than agent reasoning trajectories?
- How should verifiable process memory anchor safety-critical action logs?
- What process records would independently verify that agents performed required steps?
- Can agents themselves read and rely on tamper-evident process records?
- What evidence should benchmark operators attach to completion claims?
47 specific questions
- Should governance be applied at runtime rather than reconstructed after the fact?
- What counts as a mature governance model for agentic AI systems?
- Can runtime rules and agent loops replace pre-release governance frameworks?
- Does encoding governance into runtime loops scale as deployment environments become more complex?
- Can a single manager policy work across vastly different agent architectures?
- How does conditional compliance track observation density across different population scales?
- Can individual permissible actions collectively violate system-level constraints?
68 specific questions
- How do benchmark filtering practices hide specific model failures?
- Where do outcome grades come from once a model enters deployment?
- How should benchmarks balance verifiability against outcome resolution?
- How much do cross-model improvements compare to the original performance gains on the training benchmark?
- Why do standard accuracy metrics ignore set-level consumption constraints?
- How can a second performance metric reveal shortcuts that a single metric would hide?
- How does the absence of failure rate information affect generalizability claims?
48 specific questions
- Which evaluation habits keep safety-critical failures hidden in AI systems?
- What does it mean for errors to remain visible, contestable, and recoverable?
- Why do evaluation habits hide safety-critical challenges from view?
- How do response-centered evaluation assumptions hide safety-critical failure modes?
- What conditions allow technical systems to escape critical evaluation?
- What makes a model's errors visible and contestable to users?
- Can AI outputs inspire new directions even when they seem like failures?
53 specific questions
- What tensions emerge when AI models generate interfaces instead of rule-based systems?
- Should GUI agents use intermediate structured representations instead of raw pixels?
- Do users notice when generative interfaces don't match their own stated design principles?
- Can specialized perception components replace end-to-end vision in GUI agents?
- Should GUI perception happen inside or outside the foundation model?
- Can screen perception be effectively decoupled from planning in GUI agents?
- Why does explicit screen parsing outperform pure vision in GUI agents?
70 specific questions
- Why do stronger local checks not close the component-to-system safety gap?
- Why do sequences of safe actions sometimes violate system-level constraints?
- Why does a control blocking one moment fail against agents acting across time?
- Can outcome-only safety reports hide dependencies on server-side filtering rather than alignment?
- Why do individual safe actions create unsafe behavior collectively?
- Why does treating evaluation as a local output problem miss security risks?
- Why is evading detection easier than internalizing safety norms?