INQUIRING LINE

An AI coding assistant edits its own instructions and tools, with a human checking each change — can that process alone push it to top scores?

Can harnesses that rewrite themselves through reviewed commits achieve state-of-the-art performance?

This explores whether an agent harness (the scaffolding of prompts, tools and control code around a model) that edits its own code, with each change reviewed like a software commit, can reach top benchmark scores, and whether the self-editing is what gets it there.


This explores whether a harness that rewrites its own prompts, tools and core logic, with each change going through a review step like a code commit, can reach the top of the benchmarks, and whether the self-rewriting is the reason it does. The short answer is that it has posted top scores, but nobody has yet shown that the self-rewriting caused them. The clearest case is Ouroboros, which evolves itself through reviewed commits and reports 86.97% on Terminal-Bench 2.1 and 90.69% on OSWorld-Verified Does self-editing through reviewed commits improve agent performance?. The paper has no ablation, though: it never compares the evolved harness against the same harness frozen at its starting point. So we can't tell how much of the result comes from evolution and how much from a well-built starting harness running on a strong model.

This gap matters because harnesses that don't rewrite themselves already reach similar numbers. StateM tunes the execution system around frozen model weights and reaches 95.3% on Terminal-Bench 2.1 with GPT-5.6. Its same runbook also transfers unchanged to newer models Can execution harnesses lift model performance without retuning weights?. A large-scale automated search across 51 tasks found four reusable mechanisms (how actions run, how context is compressed, how observations are handled, and handing reading off to helpers). Together they cut token use by almost half without losing accuracy Can agent harnesses be automatically optimized across many environments?. The lesson is that harness engineering alone is a large lever. A self-rewriting harness that scores well might be pulling that lever through evolution, or it might simply have started with a good design.

The closer studies are also skeptical about what self-edits actually learn. When researchers inspect evolved harnesses, most edits look sensible but mostly save task-specific fixes the agent could have found again in a single attempt. They store answers that were already within reach instead of turning real failures into successes Do harness edits learn reusable strategies or memorize task fixes?. Another finding is less intuitive: models at every capability tier are about equally good at writing useful harness edits, but how much they gain from those edits peaks at mid-tier models. Weak models fail to use the harness at all, and strong models don't follow its instructions faithfully Do stronger models always evolve harnesses better?. Models also struggle to keep their useful intermediate updates as a harness evolves, and a harness's quality can change sharply depending on which model runs it Can language models build and maintain their own agent harnesses?.

Other work points to designs that make self-improvement more trustworthy. ModularRSI evolves each harness component separately, on data kept apart from the benchmark, and pools evidence across tasks before changing anything. That design is aimed squarely at the memorization problem, and the resulting gains hold up on unseen tasks Can harness modules improve separately from benchmark data?. The Darwin Gödel Machine is an earlier example of the same idea. It keeps an archive of agent variants, tests each one empirically, and found real capabilities such as better code editing, with roughly 2.5× gains on SWE-bench Can AI systems improve themselves through trial and error?. Ouroboros's reviewed commits fit what autoresearch work says a domain needs: version control, fast iteration, a modular design and an immediate numeric score What makes a research domain suitable for autonomous optimization?.

The bigger picture is why this question matters. Lilian Weng argues that the near-term route to recursive self-improvement, meaning AI systems that improve themselves, runs through harness engineering, not models rewriting their own weights Does recursive self-improvement start with harness engineering?. If that's right, self-rewriting harnesses are an early version of the whole project. Even so, the evidence so far shows that the scores are achievable but not that self-rewriting is what produces them. The useful thing to look for in this kind of paper is a comparison against a frozen copy of the same harness, which is exactly what is missing so far.


Sources 10 notes

Does self-editing through reviewed commits improve agent performance?

Ouroboros, a harness that rewrites its tools, prompts, and core through reviewed commits, achieved 86.97% on Terminal-Bench 2.1, 90.69% on OSWorld-Verified, and 0.2301 on CL-Bench. However, the paper lacks ablation studies to isolate evolution's actual contribution to these scores.

Can execution harnesses lift model performance without retuning weights?

StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.

Can agent harnesses be automatically optimized across many environments?

Cross-environment harness optimization yielded four mechanisms (action execution, context compaction, observation handling, delegated reading) that reduced token traffic by 44.7–49.0% while maintaining comparable performance on a 51-task benchmark, suggesting harness-level gains are orthogonal to model improvements.

Do harness edits learn reusable strategies or memorize task fixes?

Analysis of evolved harness trajectories shows rational, well-motivated edits across prompt and tool layers, but most persist fixes an agent could rediscover in a single rollout. Gains remain limited because memorized shortcuts cache what's already within reach rather than converting hard failures into successes.

Do stronger models always evolve harnesses better?

Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.

Show all 10 sources
Can language models build and maintain their own agent harnesses?

Research shows that LLMs vary sharply in building harnesses across domains, struggle to retain useful intermediate updates during evolution, and produce harnesses whose performance shifts dramatically with different executors—demonstrating that harness quality cannot be inferred from downstream task scores alone.

Can harness modules improve separately from benchmark data?

ModularRSI evolves harness modules independently using contrastive trajectories on benchmark-disjoint data, showing consistent gains across unseen tasks and domains. The approach isolates mechanism-level improvements from task-specific adaptation by aggregating evidence across tasks before updating components.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

What makes a research domain suitable for autonomous optimization?

Autonomous research pipelines require immediate scalar metrics, modular architecture, fast iteration cycles, and version control. Domains lacking any property resist autoresearch regardless of LLM capability, because the bottleneck is environmental structure, not model power.

Does recursive self-improvement start with harness engineering?

Weng argues RSI's initial path moves through optimizing deployment harnesses—instruction prompts to harness code to optimizer code—rather than models directly rewriting weights. This staged progression mirrors how prompt engineering gave way to instruction tuning while interface needs persisted.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.