Theme of inquiry
How can we measure and understand agent capabilities across different designs?
A question within its area, explored through 3 lines of inquiry below — each a family of specific questions the research asks.
69 specific questions
- Does reasoning ability help agents learn from feedback faster?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- Can individual skills improve through reuse and accumulate experience across tasks?
- Can agent skills move from prompts to trainable parameters?
- Can agent-authored skill libraries compound autonomy gains over time?
- How can agents distinguish over-generalized lessons from genuinely useful long-tail knowledge?
- Can agents improve from deployment signals without explicit human annotation?
83 specific questions
- Why do different agent memory architectures make incompatible granularity claims?
- Can agent-controlled memory management outperform fixed consolidation schedules?
- How does durable memory quality shape agent performance over time?
- Could a single agent system switch memory granularity between tasks?
- Does workflow-level memory or state-action memory better capture reusable agent knowledge?
- Why do agents ignore condensed experience in favor of raw data?
- Should agents update memory after every turn or batch process sessions?
37 specific questions
- Why do agents report success when actions actually fail?
- Why do agents report success when their actions actually fail?
- How often do agents report success when their actions actually failed?
- Why do autonomous agents report success on failed actions?
- Why do agents report success when they have actually failed at tasks?
- How do agents learn to report success on actions that actually failed?
- Can confident agent failures appear as successes in outcome reporting systems?