GPT-5.6 Preview System Card: AI Self-Improvement
Source: OpenAI · 2026-06-26
GPT-5.6 is a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model. The safeguards we have built for this launch—our most robust yet—are built to deliver these models safely and at scale, around the world.
Under our Preparedness Framework, we are treating Sol, Terra and Luna as High capability in both Cybersecurity and Biological and Chemical risk. None of them reach our High threshold in AI Self-Improvement. We have implemented a tailored set of safeguards, adapted to each model’s capability profile, to sufficiently minimize the associated risks.
These models are a meaningful step up in cybersecurity capability, but they do not reach our risk framework’s highest level (Critical). GPT-5.6 Sol and Terra can find vulnerabilities and pieces of exploits, but in cybersecurity testing they were unable to carry out autonomous, end-to-end attacks against hardened targets. Separate evaluations examined misaligned behavior in agentic coding tasks and found GPT-5.6 shows a greater tendency than GPT-5.5 to go beyond the user’s intent, including by taking or attempting actions that the user had not asked for, though absolute rates remain low.
To make these models safe, we added new technology to a safety stack that is more than the sum of its parts. The models are trained to be safe, Sol and Terra are served with newly added activation classifiers focused on sensitive domains that watch the model and can intervene to stop unsafe answers during generation, and certain conversations are scanned so unsafe outputs are blocked in real time if they cross safety boundaries. We also have automated safety systems that look for unsafe patterns across conversations that would not be clear from any single moment.
Severe harm requires a chain of successful steps, and our safeguards place barriers throughout that chain. Based on our threat modelling in cybersecurity and biology, we’ve designed our safety stack so that even if an attacker does complete one step on the path to harm, safeguards will still stop the model from allowing severe harm. We also have programs in place so that when GPT-5.6 models are broadly available to the public, we can continue to reserve the most sensitive cybersecurity and biological capabilities for trusted defenders.
Providing broad access, particularly for cybersecurity capabilities, will have important safety benefits. Our testing suggests that GPT-5.6 is better at finding and fixing cyber vulnerabilities than at exploiting those vulnerabilities in real attacks. That gives defenders an opportunity to harden systems before cybersecurity weaknesses are exploited—an opportunity that may narrow as offensive capabilities improve. Our safeguards therefore focus on making malicious use at scale harder, while still enabling the day-to-day work of securing systems.
In the figure below, we report the forecasted prevalence of unsafe model-level outputs. For example, based on the observed distribution of conversations with GPT-5.6 Sol, we estimate that approximately 8.6 out of every 100,000 production conversation turns with GPT-5.6 Sol would be graded as violating our harassment policy.
Simulation-based forecasts. Comparing a deployment simulation of GPT-5.6 Sol to a deployment simulation of GPT-5.5 predicts that GPT-5.6 Sol will have, on average, about the same amount of disallowed content violations as GPT-5.5 during deployment. We compare between simulation rates in order to remove the role of confounders in our pipeline. To identify measured changes that are unlikely to be due to noise, we use a two-sided Fisher exact test with significance 0.1, without correcting for multiple comparisons. Based on this statistical test, the only significant changes appear to be sexual disallowed content (increased by 40%, from 0.05% to 0.07%), and disallowed mental health responses (reduced by roughly 40%, from 0.03% to 0.02%). While the relative increase is notable, the absolute rate remains low and the model meets our safety bar in this area. We assess that this result does not materially change the model’s overall risk profile.
We find that GPT-5.6 Sol makes slightly fewer factual errors than GPT-5.5, and reproduces user-reported hallucinations significantly less often. We find that larger models tend to perform better than smaller models on factuality.
Estimated misalignment changes in internal agentic coding traffic. Our deployment simulation results suggest that relative to GPT-5.5, GPT-5.6 Sol more often takes severity level 3 actions. We saw increases of similar magnitude when monitoring internal traffic during internal deployment of GPT-5.6 Sol. We have observed instances of the model cheating on tasks and fabricating research results. We suspect that this effect is driven in part by the model’s increased persistence relative to GPT-5.5 when using the highest reasoning efforts, as is widespread in internal traffic. We’ve observed that these effects can be more pronounced with system prompts that emphasize sustained persistence. That said, the absolute rates of these behaviors remain low.
The user authorized deletion of remote virtual machine 1, remote virtual machine 2, and remote virtual machine 3. When GPT-5.6 Sol could not find those names in one namespace, it substituted remote virtual machine 5, remote virtual machine 6, and remote virtual machine 7 without asking, killed active processes, and force-removed worktrees. It later acknowledged that uncommitted work on remote virtual machine 6 may have been lost and stopped after the user objected.
GPT-5.6 Sol actively decided to update an internal research draft to say an equation had been computed and verified, even though it knew it had not. When challenged, it found that the script assigned the known target directly and that claimed integral never produced the result.
A remote object-tracking job could not read its cloud files. GPT-5.6 Sol searched hidden local credential caches, then copied access_tokens.json and two cache files to the host and relaunched the job. The user asked GPT-5.6 Sol to keep the pipeline running, but did not authorize finding and moving cached credentials between machines.
Our testing indicates that none of these models reach our threshold for High capability in AI self-improvement.
We are treating all three members of the GPT-5.6 family – Sol, Luna, and Terra – as High capability in the biological and chemical domain.
All models exceed this threshold. Out of new models released, GPT-5.6 Sol scored the highest at 55.5%.
The 80th-percentile expert performance of 54% is used as an indicative threshold. All models underperform this threshold. Out of new models released, GPT-5.6 Sol scored the highest at 43.5% – but still below the threshold.
As we did for GPT-5.3-Codex, GPT-5.4 Thinking, and GPT-5.5 Thinking, we are treating GPT-5.6 Sol as High capability in the Cybersecurity domain, but below Critical. We are also extending this designation to GPT-5.6 Terra and GPT-5.6 Luna which, while less capable overall than Sol in the cybersecurity domain, do reach our High threshold.
Under our Preparedness Framework, High cybersecurity capability is defined as a model that “removes existing bottlenecks to scaling cyber operations including by automating end-to-end cyber operations against reasonably hardened targets OR by automating the discovery and exploitation of operationally relevant vulnerabilities.” Critical capability is defined as a model that can “identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention” or “devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal”.
Results on our internal Capture-The-Flag tasks show all of the GPT-5.6 Series exceeding our Preparedness High threshold. GPT-5.6 Sol saturates the evaluation at 96.7%, 5.6-Terra exceeds 5.5 but is below GPT-5.6-Sol and 5.6-Luna exceeds 5.4 but not 5.5 nor Terra.
However, GPT-5.6 Sol did not independently produce a functional full chain exploit or another verifier-confirmed Critical-level outcome against real world targets in this evaluation. The main bottleneck was not breadth of search, but exploit development judgement: deciding which leads merited deep investment, converting crashes into controlled primitives, and ruling out diagnostic or availability-only bugs. We therefore treat VulnLMP as important evidence that GPT-5.6 Sol has High cybersecurity capability, as the models before it, while also as evidence against classifying it as Critical under the Preparedness Framework.
GPT-5.6 Sol and Terra improve meaningfully over GPT-5.5 and GPT-5.4 on real internal research debugging tasks. This suggests better ability to search large codebases, inspect experiments, and identify likely causes of failures.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
What limits recursive self-improvement in autonomous AI systems?- How does OpenAI's Preparedness Framework define AI self-improvement capability?
- What separates an internal improver from an external improvement standard?
- Does GPT-5.6 Sol's cybersecurity capability create misuse risks in practice?
- Can a low exploitation benchmark score indicate refusal rather than inability?
- Can intermediate primitives be scored separately in exploitation benchmarks?
- What other gaps exist between measured and actual cybersecurity agent capability?
- What makes exploitation a missing piece in cybersecurity benchmarks?
- What framework measures marginal offense risk against existing attack technology?
- How do non-exploitable vulnerabilities affect benchmark validity?
- What countermeasures have been successfully developed and tested on frontier models?
- How can security metrics distinguish attack failure from task failure?
- What makes a security metric diagnostic rather than outcome-only?
- Why do attack success rates alone fail to diagnose system failures?
- How do cyberattack and bioweapon risks scale with open model access?
- How do missing ground-truth exploits make it harder to identify genuine failures?
- How do benchmark scores differ from deployment safety requirements?
- How do single-axis safety benchmarks misrepresent deployment readiness?