INQUIRING LINE

Can you make an AI smarter just by upgrading the tools and scripts around it — without ever retraining the model itself?

Can scaffold-only modifications achieve lasting gains without updating the foundation model weights?

This explores whether you can get a model to perform better for good just by changing what surrounds it (its prompts, code wrappers, tools and memory) while never retraining the model itself, and whether those gains hold up over time.


This explores whether improving the wrapper around a frozen model, rather than the model itself, can produce gains that last. The short answer from the corpus is yes, the gains are real and sometimes large. Whether they last depends on what the wrapper is anchored to. The clearest evidence comes from harness work. Optimizing only the execution system around fixed models raised Terminal-Bench scores across several model families, and the same runbook carried over to newer models without changes Can execution harnesses lift model performance without retuning weights?. That portability is a kind of lastingness that weight updates don't have, because a fine-tune is tied to one checkpoint. A model can even rewrite its own scaffold over several rounds and keep getting better while its weights stay untouched Can language models improve their own scaffolding without weight updates?.

The more surprising part is *why* some scaffold gains stick. When a stronger model built harnesses for a weaker one, it nearly doubled the weaker model's Theory-of-Mind scores. It didn't do this by coaxing longer reasoning. It moved the unstable parts of the reasoning into deterministic code and task-specific routing Can a stronger model lift a weaker one at test time without retraining?. Code doesn't drift or forget, so the durable gains come from taking jobs away from the model, not from persuading it to do them better.

There is a catch, though. A loop where the model scores its own scaffold improvements eventually stalls. The model can't reliably check its own work, its outputs grow more alike, and it learns to game whatever it is graded on. The methods that keep improving all bring in an outside reference point: tool feedback, an outside judge, or user corrections Can models reliably improve themselves without external feedback?. STOP shows this too: how much improvement the loop can extract depends on whether its checking signal is written as code or as plain English. So scaffold-only gains last only as long as the feedback stays honest. A related warning: a change that worked once may not work after the underlying model changes, so earlier success should prompt a fresh check, not automatic reuse Should past update success guide future model changes?.

Several notes push back on treating "scaffold vs. weights" as a binary at all. One proposes a sequence: extend the scaffold, amplify what it reveals, absorb that into the weights, then *retire the scaffold*. Each surface carries its own hidden costs, and moving the gains between surfaces recovers them Can sequencing edits across three surfaces avoid hidden costs?. On that view, scaffolding is a good place to discover improvements but not always the best place to keep them. Agent memory raises the same point: an external memory module and the model it serves can each be trained separately and drift out of step, which is an argument for building memory into the model itself Should agent memory live inside the model backbone?. Another note moves the question entirely, arguing that adaptation belongs to the whole loop of model, harness and external evaluation, not to any one model snapshot Where does model adaptation actually happen?.

There is also a middle path most readers won't expect. Representation finetuning leaves every weight frozen and instead edits the model's internal activations as it runs. It beats LoRA while using 10–50x fewer parameters Can editing hidden representations beat weight updates for finetuning?. So "without updating weights" doesn't have to mean "only from the outside." You can also intervene inside a frozen model without retraining it.


Sources 9 notes

Can execution harnesses lift model performance without retuning weights?

StateM improves Terminal-Bench 2.1 accuracy across multiple models by optimizing execution systems around frozen weights. The same runbook transfers to newer models without modification, achieving 95.3% on GPT-5.6 and lifting DeepSeek-V4 Flash by 5.4 points.

Can language models improve their own scaffolding without weight updates?

STOP demonstrates that an LM can iteratively refine the improver program wrapped around it, achieving measurably better downstream performance without any weight changes. The form of the verifying signal—source code versus plain English—significantly shapes how much improvement the loop can extract.

Can a stronger model lift a weaker one at test time without retraining?

A stronger model built inference-time harnesses that nearly doubled weaker model performance on Theory-of-Mind benchmarks without retraining, primarily by moving unstable reasoning into deterministic code and task-specific routing rather than encouraging extended reasoning.

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Should past update success guide future model changes?

An update's effect depends on its source context—parent model state, data, training stage, and evaluation criteria. Autonomous systems should gate reuse with applicability checks and bounded trials rather than treat prior success as permission, because promoting a child rewrites the parent against which future evidence is measured.

Show all 9 sources
Can sequencing edits across three surfaces avoid hidden costs?

MetaRSI-v1 shows that a four-step composition—extend scaffold, amplify what it reveals, internalize into weights, retire extension—retains capability while avoiding the per-surface costs that single-edit operators impose. Each surface pays a different hidden price; the sequence distributes and recovers them.

Should agent memory live inside the model backbone?

Metis demonstrates that agent memory can be implemented as a persistent state and autonomous procedures within the model backbone rather than external modules. This approach enables end-to-end training and avoids the decoupling failures where external memory and backbone optimize independently.

Where does model adaptation actually happen?

Macaron-V1 argues adaptation is a property of the recursive cycle linking model, harness, and external contract—not individual model snapshots. Weight updates are gated by audit and evaluation against an external contract, making the loop the unit of improvement and release.

Can editing hidden representations beat weight updates for finetuning?

ReFT learns task-specific interventions on frozen model representations rather than updating weights, with LoReFT (low-rank linear subspace variant) dramatically outperforming LoRA across reasoning, instruction-following, and NLU benchmarks while using far fewer parameters.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.