SYNTHESIS NOTE
Topics›Frontier AI Risk & RSI›this note

Does code LLM self-review prevent recursive training collapse?

When code models review their own generated outputs across multiple training rounds, can self-scoring or perplexity filters maintain quality, or do they eventually rubber-stamp degraded code? Understanding self-gate failure modes matters for safe recursive training.

Synthesis note · 2026-10-08 · sourced from Frontier AI Risk & RSI

Across four code LLMs (SantaCoder, Qwen2.5-Coder, StarCoder2-3B, Code Llama-7B) and five rounds of recursive fine-tuning on their own generated code, the paper compares three review regimes: no review, "Human-gate" review using model-independent signals (compilation, static quality rules), and "AI-self-gate" review using the code LLM's own perplexity or binary self-scoring. The result: "no review collapses fastest, Human-gate filters slow but do not stop collapse, and AI-self-gate filters can look strong early but later lose their filtering effect." In the clearest case, "the binary self-gate enters a rubber-stamp regime where acceptance scores rise while benchmark correctness falls" — on MBPP+, the perplexity filter's pass rate rises from 0.167 at round 1 to 0.235 at round 5 even as the underlying model degrades.

The paper's mechanism is that a Human gate uses an acceptance function rH(x,c) that is fixed and "independent of t and θt" — exogenous to the generator — while an AI self-gate's acceptance score rφt is produced by the same model family being retrained, so it "can drift as the generator changes." They prove AI self-gating "degenerates to ungated self-training under a self-confirming acceptance condition": once the generating model has already drifted, its own perplexity or self-score becomes miscalibrated in the same direction, so the filter stops filtering. A spectral analysis of representation covariance shows the leading variance direction gets amplified round over round relative to the rest, concentrating the output distribution regardless of which gate is used — gates only change the rate, not the eventual "collapse attractor" all four models converge toward.

This sharpens What limits how much models can improve themselves?: that framework treats the verifier as a fixed quantity whose gap with generation sets the ceiling, but here the "verifier" is the generator's own scoring head, so the gap itself erodes over iterations rather than holding constant. It also qualifies Can models reliably improve themselves without external feedback? — external anchoring is the right direction, but this paper shows a weak exogenous anchor (compilation, static rules) is not sufficient either: Human-gate filtering "preserve[s] useful validity signals but cannot stop long-horizon semantic drift." Unlike How quickly do errors compound during model self-training?, which concerns a fully unverified loop, this paper's finding is that even reviewed loops collapse once the reviewer shares the generator's distribution.

The experiments stop at five rounds, Python benchmarks, and models up to 7B parameters (with additional StarCoder trajectories only in an appendix), and the paper states its Human-gate filters are "simplified proxies for full PR review" that "do not fully check semantics, security, or design quality" — so the result does not establish that stronger, more human-like PR review would also fail, only that compile/static checks and self-scoring both do. The paper's own conclusion is narrower and more actionable than a general indictment of self-review: "stable recursive code LLM training requires exogenous verification rather than model-coupled self-review," which licenses caution specifically about AI-code-reviewing-AI-code pipelines, not about human review generally.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do models learn from self-generated outputs without cascading failures?

Related concepts in this collection 5

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 139 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

AI self-gate review of generated code degenerates into rubber-stamping while human-gate filters only slow collapse