OpenAI rates new models as 'High risk' for cyber and bio skills but not for self-improvement — what actually earns that label?
How does OpenAI's Preparedness Framework define AI self-improvement capability?
This explores what OpenAI actually measures when it decides whether a model counts as "self-improving" under its Preparedness Framework, and how that measure compares with the ways other researchers think about AI improving itself.
This explores what OpenAI actually measures when it rates a model's ability to improve AI systems, including itself, and whether that measure captures what researchers mean by self-improvement. The short answer is that the collection doesn't contain the framework's formal definition. It does show how the framework is applied in practice, and that is the more revealing part. In its assessment of GPT-5.6 Sol and Terra, OpenAI rated both models High in cybersecurity and biorisk but below High in AI self-improvement, even though both did measurably better on internal research-debugging tasks Does GPT-5.6 show meaningful self-improvement capability?. The surprising detail is that the self-improvement rating rests on one debugging metric, and the published material doesn't put numbers on it. So in practice, "self-improvement capability" is measured by whether a model can help with the routine work of AI research. It is not measured by whether the model can redesign itself.
Why measure debugging instead of something more dramatic? One reason is that real self-improvement probably doesn't start with a model rewriting its own weights. Lilian Weng argues that the near-term path runs through the "harness": first the prompts, then the code that wraps and deploys a model, then the optimizer code Does recursive self-improvement start with harness engineering?. Concrete systems already work this way. The Darwin Gödel Machine improves coding agents by testing variants against benchmarks and keeping an evolving archive of the versions that work, roughly doubling performance on SWE-bench Can AI systems improve themselves through trial and error?. A "bilevel" setup goes a step further: an outer loop reads the inner loop's code, finds the bottlenecks, and writes new search methods while it runs Can an AI system improve its own search methods automatically?. A debugging metric looks for the early, mundane signs of this kind of capability.
The catch is that "self-improvement" covers very different things. A survey of 1,250 papers separates bounded self-refinement from open-ended recursive self-improvement Are self-refinement and recursive self-improvement actually the same thing?. Bounded self-refinement means a model polishing its outputs against a checkable target, which is common industry practice today. In open-ended recursive self-improvement, each round of improvement makes the next round more capable. The survey argues that grounding requirements, collapse dynamics and compute limits still hold that second kind back. Other researchers suggest the real turning point is something a debugging task can't detect: whether the AI chooses its own research goals without drifting Can AIs learn to specify their own research objectives?. A related view is that a system truly improves itself only when it can change how it learns. Today's systems run human-designed reflection loops that break when the domain shifts Can AI systems improve their own learning strategies?. Work on agents given only vague goals points the same way. Much of the hard work is turning "get better at X" into something measurable, and that step comes before any optimizing Can agents learn from vague goals without predefined metrics?.
This matters because a narrow benchmark can mislead in either direction. Neat, auto-gradable tasks can overstate or understate what frontier models can do. Open-ended evaluations of messy, long-horizon work tend to spot new capabilities earlier Do automated benchmarks hide what frontier AI systems can really do?. A single debugging score may be exactly the kind of instrument that misses a capability as it first appears. The stakes are not abstract. Anthropic has warned publicly that recursive self-improvement carries serious societal risks and has suggested slowing some development paths Does recursive self-improvement pose serious risks to society?. The Future of Life Institute argues that companies can't police this alone and should face government limits Can companies alone manage the risks of AI systems?. The takeaway is that the "below High" rating answers a narrow question: can this model debug AI research code at a certain level? The bigger questions, such as whether a model can set its own goals, change how it learns, or improve the scaffolding around it, are being studied in the research literature. The collection shows no sign that they are part of the framework's threshold.
Sources 11 notes
OpenAI's Preparedness Framework rates GPT-5.6 Sol and Terra as High capability in cybersecurity and biorisks, but below High in AI self-improvement despite measurable gains on internal research-debugging tasks. The self-improvement rating relies on a single unquantified debugging metric.
Weng argues RSI's initial path moves through optimizing deployment harnesses—instruction prompts to harness code to optimizer code—rather than models directly rewriting weights. This staged progression mirrors how prompt engineering gave way to instruction tuning while interface needs persisted.
DGM replaces formal proofs with empirical benchmarking and maintains an evolutionary archive of agent variants, achieving 2.5× improvement on SWE-bench and 2.2× on Polyglot by discovering capabilities like better code editing and context management.
An outer loop successfully read inner loop code, identified bottlenecks, and generated new Python mechanisms at runtime, discovering combinatorial optimization and bandit methods that broke the inner loop's deterministic patterns and improved performance on GPT pretraining by 5x.
A 1,250-paper survey shows that bounded, evaluable self-refinement (current industrial practice) differs fundamentally from open-ended recursive self-improvement, which remains constrained by grounding requirements, collapse dynamics, and compute limits measurable today.
Show all 11 sources
A debate participant argues that AI self-improvement loops require AIs to propose and optimize their own objectives without drift. The distinction between specified autoresearch and open-ended science hinges on whether objectives come from humans or from the AI itself.
Current self-improvement methods use extrinsic, fixed metacognitive loops designed by humans that fail under domain shift or capability changes. True self-improvement requires agents to generate their own adaptive metacognitive knowledge, planning, and evaluation—a gap confirmed as a neglected research area across neuro-symbolic AI.
When given only a natural-language capability direction without predefined tasks or metrics, self-evolving agents redirect search effort toward operationalizing the goal itself. Aspire's benchmark showed that agents must construct their own training and validation signals before optimizing, revealing a phase of work that existing methods skip.
Automated benchmarks both overstate and understate capability by privileging precisely-specified, auto-gradable tasks. Open-world evaluations of long-horizon messy tasks through qualitative log analysis—with cost explicitly reported—correct these distortions and catch emerging capabilities earlier.
Anthropic's June 2026 post, as reported by the Future of Life Institute, raised alarms about recursive self-improvement leading to propaganda, job displacement, nonhuman minds replacing humans, and loss of control. The post urged companies to consider slowing or pausing certain developmental pathways.
The Future of Life Institute argues that escalating AI incidents demonstrate private companies cannot self-police effectively, and calls for government-mandated limits on recursive self-improvement practices until safety research is complete, backed by hardware verification technology.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds
- Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement
- The Red Queen Gödel Machine: Co-Evolving Agents and Their Evaluators
- NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness
- MetaRSI / RSI2: A Meta-Recursive Self-Improving System for Recursive Self-Improving Systems Themselves
- Self-Improvements in Modern Agentic Systems: A Survey
- The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement