Theme of inquiry
How do systems improve effectively through feedback and learning?
A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.
37 specific questions
- Can evaluation environments themselves become security exposures during capability testing?
- How often do deployed models exploit evaluation environments to hack their scores?
- How do live human evaluations differ from ground-truth benchmarks?
- How does evaluation environment design become part of the security boundary?
- Is the evaluation environment itself part of the security boundary?
- Do success-only evaluations systematically overestimate real-world deployment readiness?
- Why do open-world evaluations reveal capabilities that static benchmarks hide?
29 specific questions
- What makes an evaluation criterion non-stationary enough to resist agent optimization?
- Can a progressively stricter evaluator act like a curriculum for improving agents?
- Can co-evolving evaluators alongside actors prevent reward hacking?
- Does fixed evaluation criteria saturate as self-improving agents improve?
- Can a static evaluator become the performance ceiling for an improving actor?
- Can an automated evaluator stay useful while an optimizer runs thousands of iterations?
- Should feedback channels be excluded from the reward path in agent evaluations?
18 specific questions
- How does single-turn optimization undermine multi-turn collaborative dynamics?
- Why do single-turn RL methods fail to generalize to multi-turn tasks?
- Does longer interaction horizon require fundamentally different evaluation approaches?
- How do complete multi-turn trajectories differ from isolated task examples?
- How does single-turn training undermine multi-turn strategic dialogue?
- How can interactive evaluation avoid replicating fragmentation problems from response-centered benchmark culture?
- How does bounded committed state prevent multi-turn agent failures better than transcript replay?
73 specific questions
- How does measurement error in capability benchmarks systematically underestimate or overestimate true ability?
- Why do estimates for task-level performance differ so much from full job automation timelines?
- Do reasoning benchmarks predict real performance in long delegated workflows?
- Can benchmark scores on verifiable tasks transfer to unseen problems outside the training domain?
- Why do AI benchmarks measure accuracy instead of reasoning quality?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- Why do benchmark scores not capture the true nature of AI systems?
37 specific questions
- Does a single benchmark score systematically misrepresent multi-axis agent capability?
- What agent evaluation dimensions beyond task success does a single number hide?
- Can a single capability score hide an agent's tendency to game evaluations?
- Can single-axis benchmarks measure across all three agent capability layers?
- What makes some agent benchmarks measure interaction quality better than others?
- Can trajectory analysis replace one-shot task success as the primary evaluation metric?
- Can single benchmarks predict whether an agent will work in the real world?
25 specific questions
- Why does greater automation actually obscure rather than eliminate research failure modes?
- What makes automated research results fail to generalize to held-out tasks?
- Where do human researchers retain competitive advantage over autoresearch systems?
- Does refining around bad results risk cascading errors in automated research?
- Can brute-force experimental volume substitute for human research intuition and taste?
- What makes evaluation tamper-proof enough for autonomous research systems?
- How do high-leverage decision points differ across research versus production tasks?