Theme of inquiry
How can we effectively evaluate AI systems and drive improvement?
A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.
31 specific questions
- What makes trajectory quality matter more than one-shot task success?
- Can trajectory analysis replace one-shot task success as the primary evaluation metric?
- Should agent evaluation include trajectory quality beyond final success?
- What trajectory-level metrics matter beyond one-shot task success?
- Should agent evaluation include trajectory quality and memory hygiene alongside task success?
- What dimensions should trajectory-level scoring capture beyond final correctness?
- Do trajectory quality metrics predict agent safety and user trust?
83 specific questions
- How does measurement error in capability benchmarks systematically underestimate or overestimate true ability?
- Can benchmark scores on verifiable tasks transfer to unseen problems outside the training domain?
- How can benchmark accuracy scores mask the absence of interpretable reasoning structure?
- How does a single score mix exploitation ability with task capability?
- Can memorization inflate apparent capability on benchmarks with available solutions?
- Why do benchmark scores not capture the true nature of AI systems?
- Do perfect accuracy scores hide broken internal representations?
44 specific questions
- Do single-step retrieval systems with sophisticated synthesis qualify as deep research?
- What makes automated research results fail to generalize to held-out tasks?
- Do gains in optimization benchmark scores translate to gains in real research efficiency?
- Why do per-turn thinking budgets matter alongside iterative retrieval depth?
- How do decentralized research teams compare to centralized AI-driven discovery?
- Does brute force experimentation substitute for research intuition and taste?
- Can brute-force experimental volume substitute for human research intuition and taste?
64 specific questions
- How does the generation-verification gap limit AI self-improvement capabilities?
- Can AI evaluation tools solve the verification problem they help create?
- Why do AI benchmarks measure accuracy instead of reasoning quality?
- How do live human evaluations differ from ground-truth benchmarks?
- How can high benchmark performance mask broken reasoning in AI systems?
- Can verification tools keep pace with AI artifact generation speed?
- How does low verifiability change what we can measure in AI work?
40 specific questions
- Can evaluation trajectories and interaction histories replace single-answer scoring?
- Why does evaluating multiple candidates work better than judging one answer?
- How can interactive evaluation avoid replicating fragmentation problems from response-centered benchmark culture?
- Why does enlarging the evaluation unit reintroduce comparability problems?
- Do current math benchmarks measure outcomes or rhetorical plausibility?
- Can the eight-dimension rubric predict which question types need decomposition?
- What role do multi-dimensional quality frameworks play in assessing arguments versus single-metric approaches?