SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Do harness fixes or heavier training drive frontier model gains?

RSIGym tested whether improving model weights through additional training or refining the evaluation harness itself produces better performance across six frontier models. Understanding which lever works matters for allocating research effort and budget.

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

RSIGym is an "Everything as a Service" research environment that gives agents reusable training, inference, rollout, evaluation, and sandbox services so they can study Data, Harness, or Joint improvement of a target model-harness pair. Its RSI-Index averages "the fraction of the remaining performance gap closed" across five benchmarks (SWE-bench Verified, Terminal-Bench 2.0, AIME, GPQA Diamond, SkillsBench). Six frontier research models ran independent Joint-track experiments under a $500 platform-service budget per benchmark run; Opus 5 scored highest at 0.4809, raising SWE-bench Verified from 17.67% to 50.33% and AIME from 31.67% to 97.78%. But the paper's sharper finding is about which lever did the work: "the more heavily trained candidate scores lower in eight of ten comparisons; the other two gains are one task each." Opus 5's expanded-data update cut Terminal-Bench from 11/30 to 9/30; Opus 5.5's cut GPQA from 86/100 to 83/100.

Harness changes moved the needle instead. "Harness improvements address recurring execution failures" — SWE harnesses check edits before patch submission, Terminal-Bench harnesses preserve shell state and detect unproductive loops, AIME/GPQA harnesses aggregate sampled answers. With base weights unchanged, Opus 5's harness fixes alone raised its Terminal-Bench development score from 6/30 to 10/30, and DeepSeek's raised SkillsBench from 0.0232 to 0.1137. Research style tracked this: Claude agents spent longer "diagnosing failures and revising the harness" (23.70–37.49 hours per five-run session) and took the top two RSI-Index spots; when "stronger fine-tuning fails to help, it reduces the update instead of continuing to increase training." GPT agents trained earlier and quit sooner (8.69–11.85 hours, training starting after a median of 11 minutes versus 29) and scored lower. DeepSeek's one clear training win — SWE development score rising from 0.344 to 0.600 under an unchanged harness — came from fixing a model that emitted malformed tool calls, a harness-interface fix as much as a scaling one.

This extends Can execution harnesses lift model performance without retuning weights? by showing the same harness-over-weights pattern recurring across six independently-run frontier models rather than one fixed system. It sits in some tension with Does harness self-improvement memorize tasks instead of learning broadly?: that note's concern is harness edits overfitting to training-distribution tasks, while RSIGym's agents saw heavier training actively regress held-out development scores within a single run, not merely fail to generalize. The Claude agents' diagnose-then-restrain pattern also matches Do frontier AI agents actually conduct novel research or just optimize? — composing known fixes (shell-state handling, answer aggregation) rather than training breakthroughs. And it belongs with Can agent harnesses be automatically optimized across many environments? as another case where harness-level intervention, not weight change, is what reliably moved scores.

The excerpt measures one budgeted research run per model, not iterated self-improvement; it says directly that "whether the resulting systems improve subsequent research cycles remains to be evaluated." The ten training comparisons are small and drawn from choices the agents themselves made about which checkpoints to submit, so the "heavier training hurts" pattern may partly reflect agents' own restraint rather than a general law. The implication the evidence supports is narrower than "training doesn't help": at this budget and time scale, harness revision was the more reliable lever for closing gaps on these five closed-ended benchmarks, and larger budgets or genuine multi-cycle recursion remain untested.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? Can code harness improvements rival direct model scaling for capability? Can base models hide emergent misalignment through alignment training?

Related concepts in this collection 7

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 77 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

RSIGym's Joint-track trials find harness revisions drove gains while heavier training reduced scores in most comparisons