Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

Paper · arXiv 2607.24300 · Published July 27, 2026
Frontier AI Risk & RSI

Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier–deployment gap. This gap refers to the discrepancy between an agent’s self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability, and how little exogenous trust is sufficient to prevent real regressions from being deployed. To address this problem, we introduce a Sealed Exogenous Acceptance Loop (SEAL). SEAL retains self-authored tests but compares each candidate with the incumbent through a fixed harness-side audit. The agent cannot author or inspect the audit, receives only accept/reject, and the whole incumbent state is retained after a clear regression. Our experiments show that this problem often appears in heuristic learning settings. These settings require trial-and-error discovery of the target objective. We further find that failures of self-written verification are stratified by capability. Weaker agents tend to damage previously acquired strategies behind easy self-tests. Stronger agents are more stable, but they still mismeasure the deployment distribution. Standard self-written constraints do not reliably close this gap. In contrast, SEAL outperforms unprotected baselines across six models and three random seeds. Reliable self-improvement need not abandon self-verification, but it requires at least one deployment-acceptance signal outside the agent’s control.

Introduction. Self-improving agents are moving from one-shot problem solvers toward continuous self-rewriting systems (Zhang et al. 2025; Song, Cai, and Zhao 2026). In these systems, models do not acquire new capabilities via gradient-based weight updates. Instead, they accumulate capability by reading, executing, and revising policies, heuristic rules, or tools (Novikov et al. 2025). The updated object is not the model weights, but an evolving software (Iacob et al. 2026). Heuristic Learning frames this form of self-improvement as a non-gradient learning paradigm (Weng 2026). Unlike fixedspecification program synthesis, heuristic tasks require repeated trial and error. The agent needs feedback to discover which edits improve the underlying policy. Verification is therefore infrastructure for the selfimprovement loop. A fixed-specification task can expose an external acceptance criterion directly, but open-ended heuristic learning often asks the agent to maintain tests, examples, and proxy metrics that preserve previously acquired behavior. This is attractive because a human cannot rewrite a complete specification after every edit (Barr et al. 2014). It also creates a structural conflict of interest: the agent can alter not only the policy being optimized, but also the measuring instrument used to judge it. Passing verification may then reflect a genuine improvement, an easier sample distribution, a more local metric, or tests that share the candidate policy’s own mistaken assumption. This problem is different from a one-off act of reward hacking. Prior work shows that optimizers can exploit fixed proxies, scoring scripts, and incomplete specifications (Khalifa et al. 2026; Thaman 2026; Gallego 2025). In these cases, the verifier remains fixed. Self-improving agents break this assumption: the verifier itself becomes part of the evolving system. Under this regime, the cheapest path to passing verification is not necessarily improving the policy. It can instead be selecting easier test distributions, lowering self-evaluation thresholds, or concentrating evaluation on local cases that do not match deployment. This does not require explicit cheating. Even purely local optimization of self-test accuracy can lead to a system where self-scores increase while real deployment performance degrades. We call the divergence between the agent-visible self-verification signal and an agent-hidden target evaluation the verifier–deployment gap. We ask a central question:

Can a self-improving agent reliably validate its own improvements?

Figure 1 illustrates the premise. The agent edits a policy and its tests together. Its self-score can therefore remain high even when the policy fails under hidden deployment conditions. The same policy may appear successful to the agent but fail at deployment, creating a verifier–deployment gap. If self-authored tests cannot distinguish genuine improvement from a shared error, asking the agent to evaluate itself more carefully may still be insufficient. This motivates a source of evidence that the agent does not author or control. We study this structure in a controlled policy-and-test co-evolution framework. At each round, a language model reads the current code, visible feedback, and self-test outcomes, then submits a candidate policy and candidate tests. We record deployment truth offline from an evaluation that never enters the agent context, and compare five acceptance mechanisms: no protection, monotone self-test constraints, an agent-authored discriminative check, an endogenous gate, and our Sealed Exogenous Acceptance Loop (SEAL). SEAL neither replaces the agent’s tests nor invokes another supervisor model. It compares the candidate and incumbent through a small, fixed set of harness-side hidden rollouts, returns only one accept/reject bit, and rolls back the entire policy–test state after a clear regression. Our experiments reveal a systematic failure mode. When verification is self-authored, agents can improve their own validation signal without improving the deployed policy. This produces a persistent verifier–real gap. The gap is also stratified by capability. Weaker agents often overwrite useful strategies while maintaining easy self-tests. Stronger agents are more stable, but they continue to mismeasure shifted deployment distributions. Standard self-written constraints do not reliably close the gap. In contrast, SEAL blocks the most damaging deployment regressions with only a small sealed external anchor. These findings suggest that the key bottleneck is not self-written testing itself, but the loss of an external acceptance boundary when verification becomes endogenous. Our main contributions are as follows: • We formalize the ability of self-improving agents to validate their own improvements as the verifier–real performance gap and show that this gap exhibits clear capability stratification.

Related work. Self-improving agents have received increasing attention in recent years, as large language models begin to form continuous improvement loops through feedback, memory, tool use, and code editing. Early works such as Self-Refine (Madaan et al. 2023) and Reflexion (Shinn et al. 2023) study how models can use self-feedback or environmental feedback to improve subsequent outputs. More recent self-improving coding agents (Robeyns, Szummer, and Aitchison 2025) push this idea into persistent software objects, where models accumulate capability by maintaining skill libraries (Wang et al. 2023; Zhang et al. 2026), editing codebases (Jin et al. 2025), or evolving programs (Romera-Paredes et al. 2024). Some studies find that intrinsic self-correction is unreliable without informative external feedback (Huang et al. 2024; Kamoi et al. 2024). Heuristic Learning (Weng 2026) frames this trend as a non-gradient learning paradigm. The updated object is no longer a set of neural network parameters, but a software system composed of code, rules, state representations, tests, logs, memory, and update mechanisms. Our setting lies within this paradigm. The agent improves by editing codelevel procedural policies and relies on tests and feedback to decide which versions should be retained. This shifts our focus from whether code-editing agents can improve to whether the verifier that accepts such improvements remains meaningful when it is also self-authored. This focus differs from existing agent benchmarks and proxy-objective failure studies. Benchmarks such as SWEbench (Jimenez et al. 2024) and BALROG (Paglieri et al. 2025) evaluate agents under externally specified protocols. Reward hacking, Goodhart’s law, specification gaming, and recent coding-agent failures show that optimizers can exploit gaps between tests, proxies, and true objectives (Huang et al. 2024; Chen et al. 2025). In contrast, we study a setting where the verifier is not merely an external target to exploit, but part of the evolving system state. We therefore treat selfauthored verification as an endogenous failure source and use SEAL as a minimal sealed intervention to test what external acceptance boundary is needed to prevent real deployment regressions.

Method. We study self-improvement as an iterative code-editing loop in which an agent can revise both a policy and its own verification tests. Table 1 summarizes the notation. We first define the loop and the verifier–deployment gap used to measure self-authored verification failure. We then introduce SEAL as a minimal exogenous acceptance loop before each candidate deployment. SEAL does not replace self-authored tests. Instead, it adds a small exogenous acceptance boundary before deployment while allowing the agent’s own verifier to continue evolving.

Verifier–Deployment Gap At the start of round t, the system stores an accepted state xt = (πt, Tt). The agent observes the current policy, selfauthored tests, previous accept/reject feedback, and all selftest outputs; together these form the visible transcript Ht. It then submits a candidate state x′ t = (π′ t, T ′ t). The system then obtains two conceptually distinct signals: (1) The agentauthored tests produce a self-report score rt = VT ′ t(π′ t). (2) A deployment evaluation produces deployment truth YD(π) = E[Score(Rollout(π; D))]. YD(π) is the true performance measured by an agent-hidden evaluation and used only for offline reporting. The self-report is generated by tests written or maintained by the agent, so it enters the agent’s future decisions. Deployment truth comes from an evaluation that the agent cannot observe and is not written into any agent-visible file or prompt. This setup preserves the self-authored verification capability required by self-improving agents. At the same time, it allows us to observe whether self-verification separates from true deployment performance. The central measurement in this paper is the verifier– deployment gap. It measures the discrepancy between the agent’s self-authored verification signal and deployment truth. The verifier–deployment gap is present when rt remains high while YD(πt) is low, decreasing, or unable to recover. To avoid hiding trajectory shapes behind a single normalized score, we report final deployment truth, peak deployment truth, regression rate, self-report, and structural classes of test edits.

Information Limit of Endogenous Evidence if TV(P+, P−) ≤δ, at least one error is at least (1 −δ)/2. When the two worlds are difficult to distinguish from agentvisible evidence, no endogenous-only gate can make both errors small. This is a conditional information limit: it does not show that endogenous evidence is always uninformative, nor that SEAL is minimal or unique. An additional signal helps only to the extent that it increases distinguishability in adeployment-relevant direction. The proof and qualifications appear in the supplemental materials.

Sealed Exogenous Acceptance Loop As shown in Figure 2, SEAL is a pre-deployment acceptance loop inside each round of self-improvement. The agent still generates a candidate policy and candidate tests, and self-authored tests remain available for debugging and future proposals; the self-report simply no longer has sole authority to deploy. SEAL adds a fixed harness-side audit A that compares the candidate and incumbent on hidden instances disjoint from training, self-tests, and the final deployment evaluation. The audit is not a supervisor language model or a human judge. It is an executable, task-native scoring procedure. At the beginning of round t, the system holds the current accepted policy πt and the current self-authored tests Tt. The agent submits (π′ t, T ′ t). The candidate tests T ′ t produce the self-report rt. SEAL does not use this self-report as the sole acceptance criterion. Instead, SEAL compares the candidate policy π′ t against the current policy πt using an audit A that is external to the agent’s self-authored verifier. Given hidden audit instances or seeds SA, audit dynamics or input distribution DA, and m repeats per instance, SEAL scores a policy as Design Conditions SEAL relies on four design conditions, as shown in Table 2. These conditions make the audit a minimal trusted core rather than another proxy that the agent can optimize. These conditions address different failure channels. Without exogeneity, the agent can still select easier distributions.

Discussion. Finding 1: Self-Authored Verification Loses Deployment Meaning Across Games The main results are shown in Table 4. Among the 35 model– game cells with valid self-test outputs, all end with a selfscore of at least 0.70. Yet 15 of the 35 completed policies score below their game’s random reference, including six policies at Pong’s -21.0 floor. High self-scores therefore coexist with failed deployment behavior across both models and games. Thus, self-authored verification does not reliably track deployment performance. These results separate two failure modes. In failure to discover, the agent has not found a useful policy even though its self-tests have saturated. In failure to retain, the agent first discovers useful behavior and later edits it away while tests evolve to share the new policy’s mistaken assumptions. The latter is especially consequential for continual self-improvement because the system loses capability it had already acquired. Figure 4 normalizes truth using the highest final mean. This experiment supports two claims. First, a near-perfect self-test pass rate does not guarantee true performance on the deployment distribution. Second, the failure mode changes with capability and discovery difficulty. Several models remain floored on difficult no-hint games while still selfreporting success. Stronger models can sometimes discover nontrivial policies, but their self-authored tests still mismeasure the shifted deployment distribution. Thus, the core phenomenon is not simply that weak models cheat and strong models do not. Rather, self-score can lose deployment meaning across the capability ladder.

Finding 2: Internal Self-Test Constraints Do Not Close the Gap Breakout ablation. Breakout provides the most complete condition ablation. Figure 3 provides a concrete Breakout case showing how the five conditions produce different deployment outcomes on a shared replay seed. The final policies under none, monotone, discriminative, and endogenous-gate score only 3, 2, 6, and 5 in one complete episode, respectively, and quickly miss after their paddle motion becomes misaligned with the ball trajectory. SEAL instead rejects the regressing update, retains an incumbent that remains aligned through the rebound, and reaches a return of 32 on the same replay seed. Table 5 reports three-seed means and half-ranges for all models. The results reveal a consistent problem: stronger-looking self-authored constraints do not ensure better deployment performance. Both monotone and discriminative fall below none for four of the six core models, while final self-scores remain between 0.78 and 1.00 across conditions. The reason is that monotone only checks whether tests were visibly weakened, and discriminative only checks whether the candidate beats a trivial baseline. An agent can therefore strengthen tests on an unrepresentative distribution, or easily beat a weak baseline, while still losing behavior already achieved by the current policy. SEAL directly supplies this comparison. Before deployment, its confidential exogenous audit compares the candidate with the current policy and rejects clear regressions. SEAL exceeds none for all core models, including 7.9→30.0 for Gemini, and exceeds all three conditions without an exogenous audit. Audit leakage. The final two columns of Table 5 differ only in whether numerical audit scores are revealed after rejection. SEAL is at least as high as leaky-anchor in all six rows and is strictly higher in five. The largest displayed reversal is Doubao-Seed-2.0-Pro, where the mean changes from 35.1 under SEAL to 12.7 with numerical audit-score disclosure. Thus, exposing audit scores provides no additional benefit and can substantially weaken the gate, and audit confidentiality is an important part of SEAL.

Conclusion. This paper studies self-authored verification in selfimproving agents. When a self-improving agent grades its ow n homework, failure does not always appear as explicit cheating. To characterize this phenomenon, we build a self-improvement framework in which policies and tests co-evolve over multiple rounds. Experiments show that selfauthored verification failure is not merely explicit cheating or an isolated artifact of a single model or domain. Weaker agents often corrupt useful strategies after partial discovery while preserving easy self-tests. Stronger agents are more stable, but can still mismeasure shifted deployment distributions. Standard self-authored constraints, such as monotone test strengthening or discriminative checks against trivial baselines, do not reliably close the gap. We further propose SEAL as a minimal sealed intervention to test what external trust boundary is needed to prevent real regressions from entering the accepted trajectory. The results suggest that reliable self-improvement requires a low-leakage exogenous acceptance signal that the agent cannot write, observe, or directly optimize. Our findings turn the question of whether agents can validate their own improvements from an implicit assumption into a measurable research problem, and show that some external acceptance boundary is necessary when verification becomes endogenous.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

How do AI systems determine and balance multiple competing objectives? Do evolved harnesses learn transferable strategies or task-specific optimization artifacts? Can AI agents improve their skills through accumulated experience and reuse? Why do autonomous agents misreport success on failed actions? How much of agent capability comes from harness versus the model itself? How does decomposing tasks into separate stages affect reasoning quality and safety? Can AI systems discover fundamental improvements to their own architectures? What limits recursive self-improvement in autonomous AI systems? Should governance of agentic AI systems be runtime or design-time? Can AI systems achieve real improvement without external human feedback? How do reward signal properties affect model reasoning and safety? How does awareness of evaluation context influence model behavior? How do multi-agent systems fail when coordination breaks down? How do agents learn to distinguish valuable feedback from noise? When do multi-agent systems improve over single frontier models? Do individually safe AI actions create unsafe outcomes in integrated systems? Can AI research automation sustain progress through accelerating feedback loops? How should systems validate code that agents generate?