OpenAI is postponing the release of its next major AI model, which is set to be named Astra . According to OpenAI, internal security tests did not rule out the possibility that Astra could reach the “Critical” level for cyber capabilities in the company’s in-house Preparedness Framework, OpenAI announced on August 7, 2026. Specifically, this refers to the ability to independently identify and exploit security vulnerabilities in well-protected systems without any human intervention.
Key Takeaways
- OpenAI is pausing parts of the development of Astra, the successor to the current GPT-5.6 model series.
- According to OpenAI, the reason is that a “Critical” level of cyber capabilities—the highest risk level in the Preparedness Framework—cannot be ruled out.
- Specifically, “Critical” means that Astra could independently develop zero-day exploits in hardened systems or devise complete attack strategies based solely on a broadly defined target.
- OpenAI is securing test environments, encrypting model weights more thoroughly, and having external organizations such as the UK AI Security Institute conduct additional reviews.
- Shortly before that, an internal version of Astra had caused a stir by solving ten decades-old math problems for about $2,000 in computing costs.
What exactly is Astra?
Astra is OpenAI’s next major model generation, set to succeed the current GPT-5.6 series. While GPT-5.6 Sol in ChatGPT has just received an update to reduce factual errors, OpenAI has so far been testing Astra exclusively internally. The model first drew attention not because of a security warning, but because of a mathematical breakthrough: On August 1, 2026, OpenAI announced that an internal version of Astra had solved ten math problems that had remained unsolved for decades, including a paper on the existence of non-sofian groups, a refutation of the Connes rigidity conjecture, and three problems from the catalog of mathematician Paul Erdős—one of which had remained unsolved for 80 years.
What makes this remarkable is that the solutions come with machine-verifiable certificates—that is, with formally verifiable proofs rather than mere text answers. The total computational cost for all ten solutions combined is said to have been around $2,000—an amount that would barely cover a coffee subscription for the entire editorial staff. However, an independent review by experts is still pending.
Why does OpenAI refer to “critical cyber capabilities”?
It is precisely this erratic performance on complex, formally verifiable tasks that is now making OpenAI cautious. In a blog post, the company writes that it “cannot rule out a ‘Critical’ capability level” after preliminary evaluations showed unusually strong performance on coding and penetration testing tasks. According to the Preparedness Framework, the “Critical” threshold in cybersecurity is considered to have been reached when a model can do one of two things: independently discover and develop functional zero-day exploits of all severity levels in many hardened, real-world systems without human intervention, or design and execute entirely novel attack strategies against well-protected targets when given only a broadly defined objective.
OpenAI put it this way in TechCrunch: The model can “independently identify and carry out cyberattacks against traditionally well-protected, real-world systems.” The phrase “does not rule out” is important: OpenAI does not say that Astra has definitely crossed this threshold, but rather that it reacts even to the suspicion, because its own set of rules requires exactly that.
The Preparedness Framework: OpenAI’s Set of Rules for Risky AI Capabilities
OpenAI first published the Preparedness Framework in December 2023 as an internal roadmap for how the company responds to dangerous capabilities in categories such as biological and chemical weapons, cyberattacks, or a model’s ability to improve itself. Version 2 followed in April 2025, featuring two clearly distinct levels:
| Risk Level | Meaning | Consequences for OpenAI |
|---|---|---|
| High | Significantly increased risk in a category, but still manageable | Additional safeguards required before every deployment |
| Critical | Highest level; according to OpenAI, this has never been reached in the cybersecurity domain before | Development of affected activities must be halted until enhanced security measures take effect |
According to OpenAI, Astra is thus the first model for which a “Critical” suspicion has arisen in the cyber sector, of all places. In June 2025, OpenAI had already exceeded a comparable threshold in the field of biology/chemistry and subsequently tightened the security requirements for the affected model.
These security measures are now in effect
Until the outstanding issue is resolved, OpenAI is visibly scaling back development on Astra. According to the company, the measures include:
- Internal Astra activities that do not meet the new security requirements have been paused.
- Tests are now only being conducted in isolated environments without network access.
- Model weights are more strongly encrypted, and access to tools is restricted.
- Automated monitoring of chain-of-thought outputs is designed to detect risky actions early on.
- External auditors, including government agencies and the UK AI Security Institute, are conducting additional testing of the model.
- External testing partners also receive specific security requirements from OpenAI for their own infrastructure.
OpenAI has not specified a firm release date for Astra. The wording in its own blog post suggests a delay of months rather than weeks.
Not an isolated incident: A security incident at Hugging Face as recently as July
This caution isn’t coming out of nowhere. As recently as July 2026, OpenAI acknowledged in a blog post a security incident at Hugging Face in which one of its own models played a role during an evaluation. Axios also reported on it. At the same time, OpenAI continues to roll out new features for end users at a rapid pace—for example, on August 6, 2026, it released the Adobe plugin for ChatGPT, which allows users to control Photoshop, Premiere, and other Creative Cloud tools directly within the chat. This illustrates the balancing act OpenAI is currently navigating: On the product side, development is moving forward rapidly, while on the research front—particularly regarding the most powerful models—the company is deliberately hitting the brakes.
Conclusion: A warning sign that extends beyond OpenAI
The fact that OpenAI is holding back Astra for now is more than just a single product delay. It is the first publicly documented instance in which a Preparedness Framework model is teetering on the “Critical” threshold for cyber capabilities, and a test case to determine whether the rules established by AI labs themselves actually hold up in a real-world emergency. For companies and security officials, it’s worth looking beyond the mere model question: While Astra addresses the capability side of AI risks, the EU’s mandatory AI labeling requirement, set to take effect in August 2026, addresses the transparency side in parallel. Both initiatives will be in effect simultaneously in 2026, and both demonstrate that the regulation of AI systems now takes place during ongoing operations rather than only after release.
Whether Astra will ultimately cross the critical threshold or whether OpenAI will give the all-clear after implementing additional safety measures remains to be seen. One thing is certain: Until then, the model that just solved ten math problems on the fly will remain on the shelf.