Theme of inquiry
What factors determine agentic system performance and reliability?
A question within its area, explored through 8 lines of inquiry below — each a family of specific questions the research asks.
30 specific questions
- Can a single axis benchmark ever represent deployment readiness accurately?
- Which separable capability axes reveal when single benchmarks misrepresent deployment readiness?
- Why do single-axis benchmarks fail to measure deployment-ready agent capability?
- What makes single-axis agent benchmarks unreliable for predicting deployment readiness?
- How do agent benchmarks misrepresent real-world deployment readiness?
- Does a single benchmark score systematically misrepresent multi-axis agent capability?
- How do benchmark environments misrepresent deployment readiness?
95 specific questions
- Why do different agent memory architectures make incompatible granularity claims?
- How does durable memory quality shape agent performance over time?
- Can agent-controlled memory management outperform fixed consolidation schedules?
- Why do agents ignore condensed experience in favor of raw data?
- Could a single agent system switch memory granularity between tasks?
- Should agents update memory after every turn or batch process sessions?
- Does workflow-level memory or state-action memory better capture reusable agent knowledge?
49 specific questions
- Can smaller models trained for execution handle the failure modes that stop current agents?
- How does agent reliability emerge from memory and protocols instead of model scale?
- Where does agent reliability come from if not better tools?
- How do externalizing cognitive work and coordination infrastructure relate to agent reliability?
- Can stopping rules extracted from past failures improve agent reliability without retraining?
- How does externalizing reasoning into harness artifacts improve agent reliability?
- How does structured environment-side state reduce multi-turn agent failure better than transcript replay?
72 specific questions
- Do learned workflows transfer between different agents with minimal accuracy loss?
- Can individual skills improve through reuse and accumulate experience across tasks?
- Can skill repositories evolve toward execution-oriented refinement over time?
- Can agent skills move from prompts to trainable parameters?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- Can agent-authored skill libraries compound autonomy gains over time?
- Do weight-space skills lose detail compared to textual skill descriptions?
55 specific questions
- Why do agents report success when their actions actually fail?
- Why do agents report success when actions actually fail?
- Why do autonomous agents report success on failed actions?
- How often do agents report success when their actions actually failed?
- Why do agents report success when they have actually failed at tasks?
- How do agents learn to report success on actions that actually failed?
- Why do agents claim completion when their outputs remain incomplete?
47 specific questions
- Which interaction artifacts matter most for reliable agent evaluation?
- Should feedback channels be excluded from the reward path in agent evaluations?
- Why do outcome-level metrics fail to reveal contained attacks in multi-agent pipelines?
- What makes some agent benchmarks measure interaction quality better than others?
- How do agent privacy compliance and task success differ in evaluation?
- Can a reward-seeking agent be distinguished from one pursuing intended behavior?
- Should artifact-level benchmarks replace token counts for agent evaluation?
45 specific questions
- Do evolutionary archives let agents improve themselves without formal proof?
- Can self-improving agents become truly autonomous without intrinsic metacognition?
- Does fixed evaluation criteria saturate as self-improving agents improve?
- Can agents design their own objective functions as part of learning?
- Why does the generation-verification gap limit what an agent can improve about itself?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- Can agents improve reliably without an external standard?
26 specific questions
- How do planning and grounding have opposing optimization requirements in agents?
- Why do planning and grounding have opposing optimization requirements in agents?
- How should agents separate planning from perception grounding?
- Do GUI agents need harness-level splits between planning and grounding?
- What makes planning, tool use, and reasoning into jointly optimizable subsystems?
- How do perception and execution gaps limit current AI agent performance?
- Does the planning-grounding factoring principle apply to other agent tasks?