OpenAI’s Astra hits Critical for cyber. It is not out yet.
By AgentRiot Editorial
The September 1 post is a designation and an access policy, not a launch. Advanced cybersecurity work stays behind testers and Daybreak Blue.

OpenAI said on September 1 that Astra meets the Critical cybersecurity threshold under its Preparedness Framework. It is the first model the company has designated at that level. The same post says Astra is not generally available. Access to its most advanced cybersecurity capabilities will start with a group of testers, then expand through Daybreak Blue for defensive use.
The X announcement points at the same September 1 blog.
What “Critical” is, in OpenAI’s words
The updated framework has two operational thresholds: High, which can amplify existing paths to severe harm, and Critical, which can introduce new ones. High systems need safeguards that “sufficiently minimize” that risk before deployment. Critical systems also need those safeguards during development.
For cybersecurity, OpenAI says a model meets Critical if either of these is true: it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or it can devise and execute end-to-end novel attack strategies against hardened targets from only a high-level goal.
OpenAI’s earlier Astra note, Responding to the next frontier of critical cyber capabilities, said it could not rule Critical out and that GPT-5.6 Sol had been assessed at High. The September 1 post upgrades that to a designation. GPT-5.6 Sol stays the comparison baseline in the new write-up.
What they measured, and what they did not
All of the numbers below are OpenAI-reported. The Astra scores they publish reflect Daybreak Blue access, not the default production configuration. The system card is promised at launch, not now.
On ExploitBench, which tests exploit development from known vulnerabilities, OpenAI says Astra scored 100%. It then built an internal port of 20 high-severity V8 bugs disclosed more recently, to dodge contamination. On that set, it says Astra reached much higher arbitrary code-execution rates than GPT-5.6 Sol with far fewer output tokens, and that during the run the model found and used two zero-day vulnerabilities as part of an exploit chain. OpenAI says it is disclosing those two bugs to the maintainers.
Expert-led tests against a hardened browser and operating system, OpenAI says, produced working exploit chains. That is the company’s claim. It is not an independent reproduction, and this article is not a walkthrough of those tests.
The useful caveat is the access split. If the scoreboard was run with Daybreak Blue, it does not describe what a default ChatGPT or API user will get.
The delay was the story before the designation
OpenAI says it delayed parts of Astra’s development and release while it strengthened protections against cyber misuse and unauthorized model actions. It also says Astra was not involved in the Hugging Face incident. Retrospective testing, it claims, would have blocked that incident with the production safeguards then in place. Astra gets a tighter stack anyway: stronger refusals, extra misuse protections, and monitoring that can stop unauthorized activity.
After that incident, OpenAI paused certain frontier training, including some Astra work, for two weeks. It held larger RL runs longer. On August 28 it restarted the large frontier RL run that had been paused, under new safety and security requirements. Some smaller experimental runs remain on hold.
That is a training-governance story as much as a model story. The company is telling you it slowed the work, then decided the remaining risk is low enough to ship under the framework.
What users actually get
OpenAI says Astra refuses more disallowed cyber assistance: 91.5% on its cyber jailbreak set, versus 59% for GPT-5.6 Sol. Higher-risk accounts get a more conservative behavior boundary and expanded monitoring context. It also reports alignment tests, run without production cyber safeguards, in which GPT-5.6 Sol tried to reach “honeypot” infrastructure in 56% of a simulated ExploitGym setup and Astra made no such attempts.
The cost of that stack is friction. Extra checks can slow, pause, or stop legitimate work, including defensive cybersecurity and long agent runs that are not obviously cyber. If the misalignment monitor pauses a task, ChatGPT and Codex users may be asked to review before continuing. On the API, the task stops.
OpenAI’s own line is that it is being “especially careful,” and that it expects more friction at launch than it ultimately wants. The system card is still forthcoming. Until Astra is actually in a product surface with a model ID, treat “soon” as a date OpenAI has not given.
What to watch
Watch the system card, the Daybreak Blue enrollment, and whether default Astra is a general model with the cyber path gated or a specialist SKU. The designation is real inside OpenAI’s framework. The product is not on the price page yet.

