What predicts success in ultra-long-horizon agent tasks?
Does an agent's initial solution quality matter more than its willingness to iterate? AUTOLAB's frontier-model benchmark suggests persistence through feedback loops may be the true differentiator.
AUTOLAB reframes what a long-horizon agent benchmark should test. Most agentic evals score either single-turn responses or short interactive trajectories; AUTOLAB instead hands the agent a correct but deliberately suboptimal baseline across 36 expert-curated tasks (system optimization, CUDA kernels, model development, puzzles) and asks it to improve the artifact within a strict wall-clock budget. The striking empirical result, across 17 frontier models, is that the dominant predictor of success is not the quality of the agent's initial attempt but its persistence — its willingness to repeatedly benchmark, edit, and incorporate noisy empirical feedback over many cycles. Most models, including proprietary ones, either terminate prematurely or exhaust their budget with minimal progress; claude-opus-4.6 is called out as a strong exception.
This is a sharper, more operational claim than "agents should iterate." It says the binding constraint is a behavioral disposition toward sustained empirical grounding, and that disposition is unevenly distributed across models that look comparable on one-shot benchmarks. It grounds How should we measure agent system performance beyond task success? with a concrete trajectory-level predictor, and it sits naturally alongside Does raw token spending actually predict agent performance? — persistence only pays if each loop returns informative, retained feedback, otherwise it is budget-burning churn, not progress.
The mechanism cuts against itself, however. Do models fail worse when their own errors fill the context? implies that more loops mean more accumulated mistakes in context, which should degrade the very iteration AUTOLAB rewards. The reconciliation is probably that persistence pays only when paired with calibrated scoring that lets the agent see whether an edit actually helped — pure persistence without trustworthy feedback would amplify error. That is why the authors single out harness design as the promising lever: the harness, not the backbone alone, decides whether long horizons compound feedback or compound noise.
Inquiring lines that read this note 78
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What trajectory-level metrics beyond task success best evaluate agent performance?- Do trajectory quality metrics predict agent safety and user trust?
- What trajectory-level metrics replace one-shot task success measurement?
- What trajectory-level metrics matter beyond one-shot task success?
- Should agent evaluation include trajectory quality beyond final success?
- What dimensions beyond task success matter for evaluating long-horizon agent trajectories?
- Should agent evaluation include trajectory quality and memory hygiene alongside task success?
- How can decision quality be automatically extracted from agent trajectories?
- What makes next-state signals from agent trajectories a reliable learning source?
- Why has agent research prioritized policy over world model development?
- Can world models simulate actionable possibilities instead of just predicting next states?
- What makes skills worth externalizing into a persistent harness?
- What feedback signals matter most during harness evolution search?
- How much realized agent capability comes from the harness versus the model?
- What role does effective feedback compute play in agent harness scaling?
- Can single-axis benchmarks measure across all three agent capability layers?
- What agent evaluation dimensions beyond task success does a single number hide?
- Do success-only evaluations systematically overestimate real-world deployment readiness?
- Does a single benchmark score systematically misrepresent multi-axis agent capability?
- Do automated benchmarks systematically distort what long-horizon agent capability actually looks like?
- Can a single leaderboard score capture multi-dimensional differences in agent performance?
- Can agent-authored skill libraries compound autonomy gains over time?
- Does intentionally varying environment properties isolate causal effects on agent performance?
- How can agent data flywheels improve task quality iteratively?
- How does effective feedback retention govern long-horizon agent reliability?
- How does cross-agent supervision expand the set of convergent initial conditions?
- Why does persistence in the feedback loop predict agent success better than initial solution quality?
- How do complexity, diversity, and real-world fidelity interact in agent training?
- Why do most frontier models terminate early on long-horizon benchmarks?
- Can empirical validation sustain long-term optimization without becoming gamed?
- Can expert-frontier exams discriminate frontier capability better than saturated benchmarks?
- Should evaluations shift toward open-world messy tasks instead of contests?
- Can automated benchmarks accurately capture progress on real-world long-horizon tasks?
- Why do AI agents struggle with novel experiments but excel at routine tasks?
- How does poor belief tracking cause agents to keep acting past the point of usefulness?
- Why do autonomous agents report success on failed actions?
- What causes delays between wrong decisions and visible consequences in long tasks?
- How does executable evaluation feedback sustain autonomous discovery at scale?
- Do frontier AI models fail in ways that preserve the appearance of competence?
- What distinguishes an error bound from a forecast of system behavior?
- Can stopping rules extracted from past failures improve agent reliability without retraining?
- Which harness dimensions most directly predict agent system reliability?
- Why do long-horizon agents fail when their models can solve individual steps?
- Why do outcome-only rewards fail to optimize long-horizon agent behavior?
- Can simple intrinsic reward signals emerge as effective drivers of complex capability in agents?
- Why do self-improving agents concentrate progress in the fast non-parametric loop?
- How do epoch boundaries preserve self-improvement guarantees across objective changes?
- What makes an evaluation criterion non-stationary enough to resist agent optimization?
- How would a parametric self-improvement loop differ from a non-parametric one?
- What makes an agent in an economic simulation self-evolving?
- Can agents improve reliably without an external standard?
- Can agents design their own objective functions as part of learning?
- Why does held-out evaluation matter for detecting agent overfitting?
- Should feedback channels be excluded from the reward path in agent evaluations?
- Can gradients extracted with agent-level supervision transfer across different benchmarks?
- What within-run behavioral dimensions reveal where long-horizon agents succeed or fail?
- What ecosystem conditions must exist for agents to function as economic participants?
- What structural features drive instrumental convergence across different agent goals?
- Which workplace tasks remain hardest for AI agents to complete autonomously?
- What task characteristics determine whether delegation can succeed?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
How should we measure agent system performance beyond task success?
Current evaluation metrics collapse agent behavior into a single success score, hiding critical information about how agents operate. What dimensions—trajectory quality, memory use, context efficiency, verification cost—should benchmarks actually measure?
grounds: supplies persistence/feedback-incorporation as a concrete trajectory-quality predictor
-
Does raw token spending actually predict agent performance?
Standard measures of agent effort—tokens, tool calls, operations—may not capture what makes inference-time scaling work. This explores what actually drives performance gains when agents spend more compute.
extends: persistence converts to progress only when each loop yields informative, retained feedback
-
Do models fail worse when their own errors fill the context?
As a model's prior mistakes accumulate in context, does subsequent accuracy degrade predictably? And can scaling or architectural changes prevent this self-contamination effect?
contradicts/qualifies: more iteration accumulates errors in context, so persistence helps only with trustworthy scoring
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks?
- Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
- RLVMR: Reinforcement Learning with Verifiable Meta-Reasoning Rewards for Robust Long-Horizon Agents
- LiveMCP-101: Stress Testing and Diagnosing MCP-enabled Agents on Challenging Queries
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks
Original note title
on ultra-long-horizon optimization the predictor of agent success is persistence in the feedback loop not the quality of the first attempt