Can wrapping the same AI model in smarter code and loops make it markedly better, without retraining it, and where does that stop working?
How does AI system design amplify model capabilities beyond the weights?
This explores how the scaffolding around a model, meaning the code, routing, loops and tools it runs inside, can make a fixed model perform better without changing its trained weights, and where that kind of gain stops.
This explores how the system wrapped around a model (its harness, control loops and inference-time tricks) can push performance past what the frozen weights deliver alone. The corpus gives a clear answer: a lot. It also gives a less obvious one, which is that these gains don't go evenly to every model. The cleanest evidence comes from execution harnesses. Optimizing only the system around fixed models raised Terminal-Bench scores across several models, and the same runbook carried over to newer models without changes, reaching 95.3% on one of them Can execution harnesses lift model performance without retuning weights?. In a more striking case, a stronger model built harnesses that nearly doubled a weaker model's Theory-of-Mind scores. The trick was not to make the weak model think longer. The harness moved the unreliable parts of its reasoning into ordinary deterministic code and sent each kind of task to its own route Can a stronger model lift a weaker one at test time without retraining?. So good system design often works by taking work away from the model, not by asking more of it.
Why does this work at all? One answer is that models already contain more ability than they show by default. Five separate techniques all draw out reasoning that is already present in base models: RL steering, critique fine-tuning, decoding changes, feature steering and RLVR. That points to elicitation, not missing capability, as the bottleneck Do base models already contain hidden reasoning ability?. Seen this way, a harness is one more elicitation tool, running at inference time instead of training time. Related ideas sit just past the edge of what usually counts as 'system design'. Looped models reuse their own layers to gain reasoning depth without adding parameters Can models learn by looping instead of growing larger?. Transformer² mixes task-specific expert vectors on the fly at inference Can models dynamically activate expert skills at inference time?. Both get more out of the same weights by changing how those weights are used.
The surprising finding is about who benefits. Models at every tier are about equally good at writing useful harness edits. But the ability to actually benefit from those edits forms an inverted U that peaks at mid-tier models Do stronger models always evolve harnesses better?. Weak models fail to call the harness properly. Strong models struggle to follow its instructions faithfully, perhaps because they lean on their own judgment. So system design is not a uniform multiplier. It does the most for models in the middle of the range.
The next step is systems that redesign themselves. The Darwin Gödel Machine keeps an archive of agent variants and tests each change against benchmarks instead of proving it correct. Along the way it found better code-editing and context-management methods on its own and roughly doubled its SWE-bench score Can AI systems improve themselves through trial and error?. Bilevel autoresearch goes one level higher: an outer loop reads the inner loop's code, spots bottlenecks and writes new search mechanisms at runtime Can an AI system improve its own search methods automatically?. In both cases the base model stays fixed, and the system around it is what improves.
One caveat runs the other way. Some limits that look like system or capability problems actually come from training. Agents rarely take initiative because training rewards each next turn in isolation, not because they lack the ability; RL can make that behavior trainable Why do AI agents fail to take initiative?. And a compact 35B model trained on long, execution-heavy tasks keeps up with much larger models at far lower cost Does model efficiency matter more than peak capability for real work?. The practical takeaway: harnesses, training and architecture are interchangeable levers. Often the cheapest gain comes from deciding which parts of a task the model shouldn't be doing at all.
Sources 10 notes
StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.
A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.
Five independent mechanisms—RL steering, critique fine-tuning, decoding changes, SAE feature steering, and RLVR—all elicit reasoning already present in base model activations. Post-training selects rather than creates reasoning; the bottleneck is elicitation, not capability acquisition.
Models that re-apply layers in recurrent depth outperform larger feedforward networks on reasoning tasks. This works because recursion enables state tracking and compositional generalization that parameter scaling alone cannot achieve, with convergence signals providing natural halting.
Transformer2 demonstrates that tuning only singular values within weight matrices produces composable expert vectors that dynamically mix at inference without interference, outperforming LoRA with fewer parameters and enabling continual specialization.
Show all 10 sources
Model capability to produce useful harness edits stays constant across tiers, but capacity to actually benefit from those edits follows an inverted U-shape, peaking in mid-tier models. Weak models fail to invoke harnesses; strong models struggle with faithful instruction-following.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.
Research shows next-turn reward optimization structurally removes initiative from models, but proactive behaviors like critical thinking and clarification-seeking are trainable (0.15% to 73.98% with RL). The core challenge is balancing proactivity with civility to avoid intrusion.
Occamy-1.0, a 35B-parameter model further trained on execution-grounded data and long-horizon trajectories, achieves competitive performance with much larger models while sitting at the low-cost knee of the Pareto frontier, suggesting that training for coordination and follow-through substitutes for raw scale in multi-step work.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Sharpening Tax in Post-Training
- Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- Evaluating Theory of Mind in Reasoning Models: Robustness over Reasoning
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?