Theme of inquiry

How can we effectively evaluate AI systems and drive improvement?

A question within its area, explored through 5 lines of inquiry below — each a family of specific questions the research asks.


What trajectory-level metrics beyond task success best evaluate agent performance?

31 specific questions

See all 31 questions in this line of inquiry
How do capability benchmark scores systematically misrepresent true model abilities?

83 specific questions

See all 83 questions in this line of inquiry
Can brute-force automated research substitute for iterative depth and human research intuition?

44 specific questions

See all 44 questions in this line of inquiry
How does the generation-verification gap limit what we can measure about AI reasoning?

64 specific questions

See all 64 questions in this line of inquiry
How does evaluation scope and dimensionality affect what we measure?

40 specific questions

See all 40 questions in this line of inquiry