INQUIRING LINE

Why does a self-improving AI get further when its rules are written as code instead of plain English?

Why does moving constraint descriptions change recursive improvement outcomes?

This explores why changing where or how a self-improving system's rules are written down (in plain English, in code, or as checks built into the loop) changes how much the system can improve itself.


This explores why changing where or how a self-improving system's rules are written down (in plain English, in code, or as checks built into the loop) changes how much the system can improve itself. The clearest direct evidence comes from STOP, where a language model repeatedly rewrites the program that wraps around it while its own weights stay frozen Can language models improve their own scaffolding without weight updates?. The researchers found that the form of the checking signal matters a lot. A constraint written as source code gave the loop more to work with than the same constraint written in plain English. One way to read this: a rule written in code can be run and checked, while a rule written in prose has to be interpreted, and the model can interpret it loosely.

Other notes in the collection point the same way. When agents are allowed to freely rewrite their own instructions, their skills drift and become unstable. SkillOpt gets better results by building the limits into the loop itself: a cap on how much can change per step, a held-out test that each edit must pass, and a buffer that keeps rejected edits as examples of what not to do Does constraining edits make skill learning more stable?. The ACE framework makes a related move. It updates its context in small structured increments instead of rewriting it wholesale, which stops useful detail from being lost over repeated passes Can context playbooks prevent knowledge loss during iteration?. The Darwin Gödel Machine goes furthest. It drops the requirement that a self-modification be formally proven to help, and instead tests each variant against benchmarks while keeping an archive of past versions. That change is what makes open-ended improvement workable at all Can AI systems improve themselves through trial and error?.

The less obvious lesson comes from work on guardrails. A rule only works if it sits somewhere that can actually see what it is about. A check that looks at one action at a time cannot even express a rule about sequences of actions, such as 'these steps are each fine, but together they cause harm.' Only a monitor that tracks history can Can stateless checks ever catch sequence-level constraint violations?. Self-improvement loops face the same problem. A rule placed in advisory text that the model reads before acting is a different kind of thing from a rule placed in a gate that judges the outcome afterward. So moving a constraint can change more than how strictly it is enforced. It can change whether the constraint is enforceable at all.

There is also a reason constraints work better outside the model's own text generation. Language models write one token at a time and cannot take back what they have already written, while constraint solving depends on throwing out partial answers that turn out to be invalid Why does autoregressive generation fail at constraint satisfaction?. Even frontier reasoning models score only about 20-23% on problems that require real backtracking Can reasoning models actually sustain long-chain reflection?. If a constraint lives only in the prompt, the model has to enforce it during generation, which is its weakest mode. Moving the constraint into code, a validator, or an archive gives the loop the ability to discard bad attempts, which the model lacks.

One caveat: STOP is the only source in the collection that directly compares the same constraint written in different forms. The other notes support the idea by analogy rather than by running that exact experiment.


Sources 7 notes

Can language models improve their own scaffolding without weight updates?

STOP demonstrates that an LM can iteratively refine the improver program wrapped around it, achieving measurably better downstream performance without any weight changes. The form of the verifying signal—source code versus plain English—significantly shapes how much improvement the loop can extract.

Does constraining edits make skill learning more stable?

SkillOpt's ablations show that adding a textual learning-rate budget, held-out validation gate, and rejected-edit buffer (retaining failed edits as negative feedback) produces more stable and generalizable skill improvement than allowing agents to freely rewrite their own instructions.

Can context playbooks prevent knowledge loss during iteration?

The ACE framework treats contexts as evolving playbooks using generation-reflection-curation loops rather than full rewrites. This prevents knowledge loss from compression and detail erosion, achieving +10.6% on agentic tasks and +8.6% on finance without labeled supervision.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Can stateless checks ever catch sequence-level constraint violations?

Per-action checks are structurally unable to state constraints that depend on prior history. Only stateful monitors tracking composed multi-party behavior can verify the behavioral envelopes that prevent individually permissible actions from collectively violating system-level safety.

Show all 7 sources
Why does autoregressive generation fail at constraint satisfaction?

The performance ceiling on constraint satisfaction problems is not a model-quality issue but an architectural limitation: autoregressive transformers cannot retract emitted tokens, while CSP solvers fundamentally depend on discarding invalid partial assignments. Symbolic solver integration works because it supplies what the architecture lacks.

Can reasoning models actually sustain long-chain reflection?

DeepSeek-R1 and o1-preview achieve only 20-23.6% exact match on 850 constraint satisfaction problems requiring genuine backtracking. This ceiling reveals that reflective reasoning fluency does not translate to actual problem-solving competence on unfamiliar instance structures.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.