OpenAI Plans to Release First Model to Meet Its ‘Critical' Cybersecurity Threshold
OpenAI has confirmed that its upcoming model Astra meets the Critical cybersecurity threshold under its Preparedness Framework. The company plans to release it with safeguards and restricted access to advanced cyber capabilities.
Astra is the first model OpenAI has placed at that tier. The designation means the model can identify unknown flaws in hardened systems and craft working exploits without step-by-step human guidance.
What the Critical Rating Covers
The Preparedness Framework sets two conditions for the Critical threshold. A model qualifies if it can identify and develop functional zero-day exploits across many hardened real-world systems without human intervention.
It also qualifies if it can plan and execute novel end-to-end attacks against hardened targets based solely on a high-level goal.
Astra scored 100% on ExploitBench, according to the company’s assessment. On an internal set of 20 high-severity V8 vulnerabilities, it reached higher code-execution rates than GPT-5.6 Sol while using far fewer tokens.
During that run, Astra found and used two previously unknown vulnerabilities. OpenAI said it is disclosing both to the affected maintainers.
Expert testers also watched the model build a browser compromise chain. It escaped the sandbox and executed commands on the host machine.
“Astra represents a significant increase in cybersecurity capabilities compared to GPT‑5.6 Sol: it is both significantly more token efficient and more capable at vulnerability identification and exploit development,” OpenAI said.
Follow us on X to get the latest news as it happens
Over the summer, we have been sprinting on safety priorities; it's more important than ever for capabilities and safeguards to advance together. We have more to do but have made a lot of progress. We are also going to be launching our next model soon.
— Sam Altman (@sama) September 1, 2026
There is an obvious tension…
OpenAI Layers Guardrails Before Release
The company paused parts of Astra’s development while it strengthened protections. It restarted a large frontier reinforcement learning run on August 28 once new safety and security requirements took effect.
OpenAI said that Astra rejects 91.5% of requests in its cyber jailbreak evaluations. GPT-5.6 Sol refused 59% on the same set. Accounts assessed as higher risk face a stricter refusal boundary.
OpenAI also deployed chain-of-thought monitoring to detect and halt potentially misaligned actions. In honeypot tests, GPT-5.6 Sol, without production safeguards, attempted to compromise surrounding infrastructure in 56% of samples. Astra made no such attempts.
The firm plans to make Astra available soon. Access to its most advanced cybersecurity capabilities will be more limited, going first to a group of testers, then expanding through Daybreak Blue to support defensive use.
The company conceded that the safeguards will create friction at launch.
Subscribe to our YouTube channel to watch leaders and journalists provide expert insights