Theme of inquiry

How do systems improve effectively through feedback and learning?

A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.


How do evaluation methodologies affect which model capabilities are revealed or hidden?

37 specific questions

See all 37 questions in this line of inquiry
How can evaluation criteria remain robust against agent gaming?

29 specific questions

See all 29 questions in this line of inquiry
Why does single-turn training fail to generalize to multi-turn tasks?

18 specific questions

See all 18 questions in this line of inquiry
Do AI capability benchmarks accurately measure reasoning ability or just surface patterns?

73 specific questions

See all 73 questions in this line of inquiry
Do single-axis benchmarks adequately measure multi-dimensional agent capability?

37 specific questions

See all 37 questions in this line of inquiry
What limitations prevent automated research from matching human research quality?

25 specific questions

See all 25 questions in this line of inquiry