SYNTHESIS NOTE
Topics›Evolution›this note

Does automated evolution match human-built agent performance?

Can an agent improved through automated loops in 8 days generalize as well as an agent refined through human-driven R&D? This tests whether autonomous design iteration reaches human-level quality on tasks outside the training set.

Synthesis note · 2026-09-24 · sourced from Evolution

The discussion's comparison: "On four held-out benchmarks spanning in- and out-of-distribution tasks, AIDE85 equals or surpasses AIDEhuman (section 3.3), a strong baseline developed through human-driven R&D (appendix B)." The evolved agent is labeled "AIDE85" in the excerpt, which does not define the label; I read it as the agent after the seven accepted rewrites, and that identification is mine.

The baseline is the point of the sentence. Seven rewrites accepted by an automated loop in 8 days are set against an agent that people improved by hand, and the paper's claim is that the loop's product is at least as good on tasks the loop did not select on. Read with Do AIDE2's improvements transfer to unseen tasks?, it says the automated route matched the human route on the held-out set.

What "equals or surpasses" does and does not say. It leaves room for ties on some benchmarks and wins on others, and the excerpt does not say which are which or by how much. It is a comparison to AIDEhuman, not to a test-time-search baseline at matched budget, so it does not stand in for the control in How should we measure gains from automatic harness evolution?. And it is a comparison with a human-driven R&D baseline, not with human–AI collaboration, so it does not test Can human-AI research teams improve faster than autonomous AI systems?: neither speed nor safety is compared in the excerpt, and the effort behind AIDEhuman (appendix B) is not given.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do evolved harness improvements generalize as reusable strategies or memorize? Why does solution diversity in test-time search improve model generalization? How can evaluation criteria remain robust against agent gaming?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 115 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

on four held-out benchmarks the agent AIDE2 evolved equals or surpasses AIDEhuman, a strong baseline developed through human-driven R&D