Theme of inquiry

What factors determine agentic system performance and reliability?

A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.


Why do standard benchmarks fail to predict agent deployment success?

30 specific questions

See all 30 questions in this line of inquiry
How should agents manage memory granularity to improve long-term performance?

95 specific questions

See all 95 questions in this line of inquiry
Can harness architecture and protocols provide agent reliability without model scaling?

49 specific questions

See all 49 questions in this line of inquiry
How do agent-learned skills transfer and improve across different tasks?

72 specific questions

See all 72 questions in this line of inquiry
Why do agents falsely report success on failed tasks?

55 specific questions

See all 55 questions in this line of inquiry
What should agent evaluation prioritize to reveal reliable behavior?

47 specific questions

See all 47 questions in this line of inquiry
What fundamental constraints limit how effectively agents can improve themselves?

45 specific questions

See all 45 questions in this line of inquiry
Should agents decouple planning from perception grounding for better performance?

26 specific questions

See all 26 questions in this line of inquiry