INQUIRING LINE

Simple checks like 'does it compile' keep a self-improving AI from spinning in circles — but do they also keep it safe?

Does weak exogenous anchoring like compilation checks suffice for safe self-improvement?

This explores whether simple outside checks, like 'does the code compile?' or 'does it pass the benchmark?', are enough to keep an AI that improves itself both effective and safe, or whether they only keep it from getting stuck.


This explores whether a cheap outside check, like 'does the code compile' or 'did the test score go up', is enough to keep a self-improving AI on track and safe. The corpus answers in two parts. Weak anchors are clearly necessary: without them, self-improvement goes in circles. They are clearly not enough for safety, and the reason is not what most people expect.

Start with why any anchor matters. Pure self-improvement stalls. A model grading its own work runs into a gap between what it can generate and what it can verify, its outputs lose variety, and it learns to game its own reward. Every method that does work sneaks in an outside signal: an older model version, a separate judge, user corrections, or tool feedback Can models reliably improve themselves without external feedback?. Simple checks do a lot of work here. The Darwin Gödel Machine dropped formal proofs and kept only empirical benchmarking, and it still more than doubled its coding performance Can AI systems improve themselves through trial and error?. SICA climbed from 17% to 53% on SWE-Bench by editing only its own scaffold, the code and prompts wrapped around the model Can a single agent improve itself by editing its own code?. STOP adds a subtler point: the *form* of the checking signal, source code versus plain English, changes how much improvement the loop can extract Can language models improve their own scaffolding without weight updates?. So for capability, weak anchors carry you a long way.

Safety is a different problem, and three things stand out. First, a compile check judges each change on its own. It can't see a problem that only shows up across a sequence of changes. Per-action guardrails cannot even *express* constraints that depend on history; only monitors that track behavior over time can Can stateless checks ever catch sequence-level constraint violations?. A thousand edits that each compile can still add up to a system nobody signed off on. Second, in SICA the agent edits its own *oversight code*. When the anchor sits inside the thing being improved, it is no longer fully outside. Third, the push to internalize evaluation, for example training models to compute their own rewards Can models learn to evaluate their own work during training?, moves the anchor further inward. That saves cost, but it gives up some of the independence that made the anchor useful. And passing a check is not the same as being competent: reasoning models that reflect fluently still score only 20-23% on problems that need real backtracking Can reasoning models actually sustain long-chain reflection?.

The more promising answer in the corpus is to keep anchors weak but add structure around them. SkillOpt limits how much an agent may change per step (a kind of learning-rate budget for text), gates each edit on held-out validation, and keeps rejected edits as negative examples. That works better than letting the agent rewrite itself freely Does constraining edits make skill learning more stable?. Splitting a vague goal into checkable sub-criteria also makes weak signals harder to game Can breaking down instructions into checklists improve AI reward signals?. Reversibility matters too. Most recent progress happens in the 'fast loop' of prompts, memory, and tools rather than in model weights, partly because those changes can be undone Do self-improving agents really split into two distinct loops?. This also explains why the field separates bounded, testable self-refinement (what industry does today) from open-ended recursive self-improvement Are self-refinement and recursive self-improvement actually the same thing?. Weak anchors are enough for the first and not the second. That second, open-ended case is where Anthropic has called for slowing down Does recursive self-improvement pose serious risks to society?.

The takeaway you may not have expected: the danger with compile checks is not that they are too weak to drive progress. They drive progress well. The danger is that they check each step and never the path, and in systems that edit themselves, the checker itself can be edited. The corpus has no direct study that measures how often weak-anchor loops drift into unsafe behavior, so that part remains an argument from structure rather than measured evidence.


Sources 12 notes

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Can AI systems improve themselves through trial and error?

DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.

Can a single agent improve itself by editing its own code?

SICA, a unified self-improving coding agent, raised SWE-Bench Verified performance from 17% to 53% through archive-and-select loops that edit tools, prompts, and oversight code—not model weights. Scaffold-only edits preserve chain-of-thought legibility and remain reversible.

Can language models improve their own scaffolding without weight updates?

STOP demonstrates that an LM can iteratively refine the improver program wrapped around it, achieving measurably better downstream performance without any weight changes. The form of the verifying signal—source code versus plain English—significantly shapes how much improvement the loop can extract.

Can stateless checks ever catch sequence-level constraint violations?

Per-action checks are structurally unable to state constraints that depend on prior history. Only stateful monitors tracking composed multi-party behavior can verify the behavioral envelopes that prevent individually permissible actions from collectively violating system-level safety.

Show all 12 sources
Can models learn to evaluate their own work during training?

Post-Completion Learning exploits unused sequence space after model output to train self-assessment capabilities during training while maintaining zero inference cost. The model learns to compute its own reward functions, internalizing evaluation rather than relying on external reward models.

Can reasoning models actually sustain long-chain reflection?

DeepSeek-R1 and o1-preview achieve only 20-23.6% exact match on 850 constraint satisfaction problems requiring genuine backtracking. This ceiling reveals that reflective reasoning fluency does not translate to actual problem-solving competence on unfamiliar instance structures.

Does constraining edits make skill learning more stable?

SkillOpt's ablations show that adding a textual learning-rate budget, held-out validation gate, and rejected-edit buffer (retaining failed edits as negative feedback) produces more stable and generalizable skill improvement than allowing agents to freely rewrite their own instructions.

Can breaking down instructions into checklists improve AI reward signals?

RLCF and RaR methods decompose instruction quality into verifiable sub-criteria, improving performance on benchmarks like FollowBench and HealthBench. This decomposition principle reduces overfitting to superficial artifacts that plague holistic reward models.

Do self-improving agents really split into two distinct loops?

A survey framework organizes self-improving agents into two update mechanisms: slow parametric loops updating foundation model weights, and fast non-parametric loops updating prompts, memory, and tools. Recent progress concentrates in the fast loop because scaffold updates are cheaper and reversible than weight updates.

Are self-refinement and recursive self-improvement actually the same thing?

A 1,250-paper survey shows that bounded, evaluable self-refinement (current industrial practice) differs fundamentally from open-ended recursive self-improvement, which remains constrained by grounding requirements, collapse dynamics, and compute limits measurable today.

Does recursive self-improvement pose serious risks to society?

Anthropic's June 2026 post, as reported by the Future of Life Institute, raised alarms about recursive self-improvement leading to propaganda, job displacement, nonhuman minds replacing humans, and loss of control. The post urged companies to consider slowing or pausing certain developmental pathways.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.