Do harness fixes or heavier training drive frontier model gains?
RSIGym tested whether improving model weights through additional training or refining the evaluation harness itself produces better performance across six frontier models. Understanding which lever works matters for allocating research effort and budget.
RSIGym is an "Everything as a Service" research environment that gives agents reusable training, inference, rollout, evaluation, and sandbox services so they can study Data, Harness, or Joint improvement of a target model-harness pair. Its RSI-Index averages "the fraction of the remaining performance gap closed" across five benchmarks (SWE-bench Verified, Terminal-Bench 2.0, AIME, GPQA Diamond, SkillsBench). Six frontier research models ran independent Joint-track experiments under a $500 platform-service budget per benchmark run; Opus 5 scored highest at 0.4809, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. But the paper's sharper finding is about which lever did the work: "the more heavily trained candidate scores lower in eight of ten comparisons; the other two gains are one task each." Opus 5's expanded-data update cut Terminal-Bench from 11/30 to 9/30; Opus 5.5's cut GPQA from 86/100 to 83/100.
Harness changes moved the needle instead. "Harness improvements address recurring execution failures" — SWE harnesses check edits before patch submission, Terminal-Bench harnesses preserve shell state and detect unproductive loops, AIME/GPQA harnesses aggregate sampled answers. With base weights unchanged, Opus 5's harness fixes alone raised its Terminal-Bench development score from 6/30 to 10/30, and DeepSeek's raised SkillsBench from 0.0232 to 0.1137. Research style tracked this: Claude agents spent longer "diagnosing failures and revising the harness" (23.70–37.49 hours per five-run session) and took the top two RSI-Index spots; when "stronger fine-tuning fails to help, it reduces the update instead of continuing to increase training." GPT agents trained earlier and quit sooner (8.69–11.85 hours, training starting after a median of 11 minutes versus 29) and scored lower. DeepSeek's one clear training win — SWE development score rising from 0.344 to 0.600 under an unchanged harness — came from fixing a model that emitted malformed tool calls, a harness-interface fix as much as a scaling one.
This extends Can execution harnesses lift model performance without retuning weights? by showing the same harness-over-weights pattern recurring across six independently-run frontier models rather than one fixed system. It sits in some tension with Does harness self-improvement memorize tasks instead of learning broadly?: that note's concern is harness edits overfitting to training-distribution tasks, while RSIGym's agents saw heavier training actively regress held-out development scores within a single run, not merely fail to generalize. The Claude agents' diagnose-then-restrain pattern also matches Do frontier AI agents actually conduct novel research or just optimize? — composing known fixes (shell-state handling, answer aggregation) rather than training breakthroughs. And it belongs with Can agent harnesses be automatically optimized across many environments? as another case where harness-level intervention, not weight change, is what reliably moved scores.
The excerpt measures one budgeted research run per model, not iterated self-improvement; it says directly that "whether the resulting systems improve subsequent research cycles remains to be evaluated." The ten training comparisons are small and drawn from choices the agents themselves made about which checkpoints to submit, so the "heavier training hurts" pattern may partly reflect agents' own restraint rather than a general law. The implication the evidence supports is narrower than "training doesn't help": at this budget and time scale, harness revision was the more reliable lever for closing gaps on these five closed-ended benchmarks, and larger budgets or genuine multi-cycle recursion remain untested.
Inquiring lines that read this note 3
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? Can code harness improvements rival direct model scaling for capability? Can base models hide emergent misalignment through alignment training?Related concepts in this collection 7
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can execution harnesses lift model performance without retuning weights?
Explores whether improving an agent's runtime system—not the model itself—can boost benchmark accuracy and transfer across different model versions without modification.
RSIGym's six-model comparison shows the same harness-over-weights pattern recurring across independent runs, not one fixed system.
-
Does harness self-improvement memorize tasks instead of learning broadly?
When agents automatically edit their own prompts and tools based on task feedback, do those improvements generalize to new domains or just fit the training tasks? This matters because overfitting at the harness level could hide real capability gains.
Contrasts: here heavier training regresses scores within a single run, not merely overfits out of distribution.
-
Do frontier AI agents actually conduct novel research or just optimize?
Exploring whether current long-horizon research agents generate genuine methodological novelty or primarily recombine established techniques. This matters for understanding how close we are to recursive self-improvement through AI.
Claude's diagnose-then-restrain behavior matches this engineering-optimizer posture rather than novel training breakthroughs.
-
Can agent harnesses be automatically optimized across many environments?
Explores whether scaling auto-research loops across diverse harness environments can discover mechanisms that reduce token use without sacrificing task performance, and whether such discoveries generalize.
Same family of finding: harness-level intervention moves benchmark scores more reliably than weight changes.
-
Does training editors on real outcomes beat prompting larger models?
Can a small model trained on whether its patches actually work outperform larger frontier models prompted to make the same edits? This matters because it tests whether feedback beats raw capacity for runtime system modification.
Evidence for A: Harness-R1's RL-trained engineer beats prompting larger fixed models, reinforcing that harness gains outperform heavier training
-
Does recursive self-improvement start with harness engineering?
Explores whether near-term RSI advances through optimizing deployment systems and orchestration layers rather than models directly rewriting their own weights, and what evidence supports this pathway.
Extends A: Weng predicts near-term RSI runs through harness engineering before weight rewriting, framing why harness fixes drove A's gains
-
Can harness modules improve separately from benchmark data?
Does evolving harness components independently on out-of-distribution data, using contrasted success and failure trajectories, help distinguish reusable improvements from task-specific overfitting? This matters because current methods conflate general gains with benchmark adaptation.
Extends A: ModularRSI's modular, benchmark-disjoint harness updates aim to generalize the harness-driven gains A measured across models
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- RSIGym: A Flexible Environment for Recursive Self-Improvement
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories
- Scaling Laws for Agent Harnesses via Effective Feedback Compute
- Sharpening Tax in Post-Training
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
Original note title
RSIGym's Joint-track trials find harness revisions drove gains while heavier training reduced scores in most comparisons