Line of inquiry
Inquiring lines›How do we evaluate and improve AI…›How can we effectively evaluate AI…›this line of inquiry
What trajectory-level metrics beyond task success best evaluate agent performance?
A broader line of inquiry — a family of 31 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 31
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- What makes trajectory quality matter more than one-shot task success?
- Can trajectory analysis replace one-shot task success as the primary evaluation metric?
- Should agent evaluation include trajectory quality beyond final success?
- What trajectory-level metrics matter beyond one-shot task success?
- Should agent evaluation include trajectory quality and memory hygiene alongside task success?
- What dimensions should trajectory-level scoring capture beyond final correctness?
- Do trajectory quality metrics predict agent safety and user trust?
- What dimensions beyond task success matter for evaluating long-horizon agent trajectories?
- How do execution traces reveal error propagation in multi-step agent decisions?
- What trajectory-level metrics replace one-shot task success measurement?
- How can decision quality be automatically extracted from agent trajectories?
- Can trajectory structure alone reveal process quality without human annotation?
- What makes next-state signals from agent trajectories a reliable learning source?
- Why do sparse outcome rewards fail to credit correct tool calls in failed trajectories?
- Can graph topology represent successful trajectory clusters more effectively than skill libraries?
- What makes a trajectory score interpretable across different interactive benchmarks?
- Can tool-call advantage attribution distinguish between correct and incorrect calls in mixed trajectories?
- Can parallel trajectories reveal better decision branches without human labeling?
- How do execution trajectories become valid training examples after validation?
- What hidden signals in agent logs reveal about frontier capability beyond pass-fail outcomes?
- Can influence estimation identify the most valuable trajectories in agentic training?
- Can trajectory-level visibility separate refusals from real skill gaps?
- How do we measure progress without confusing it with task completion?
- What makes trajectory more actionable than absolute scores for human moderators?
- What deployment modes work best for trajectory-aware reward signals?
- How does trajectory filtering handle noise when language models use code execution tools?
- How do chunk-based step segmentation and trajectory structure modeling differ?
- Does decision-making taste predict end-to-end task success independently?
- How do trajectory quality and memory hygiene differ as evaluation metrics?
- When is information-flow tracking worth its cost over classification?
- How does evaluating interaction trajectories change what we measure beyond correctness?