Astra is OpenAI's first 'Critical' cybersecurity model: two zero-days found in testing, access to be limited
OpenAI says Astra is the first model to meet the 'Critical' cybersecurity threshold of its Preparedness Framework, after finding two zero-days in testing. Advanced cyber access starts with a small tester group.
In a post titled Path to Astra, OpenAI states that Astra meets the Critical cybersecurity threshold of its Preparedness Framework, the first model ever placed at that level. During evaluation, it says, the model found and used two previously unknown vulnerabilities. Release is planned “soon”, with advanced cyber capabilities initially restricted to a small group.
On 30 August we covered the leaks claiming a launch “next week”; the safety brake OpenAI confirmed then is now formal, and explained.
Zero-days and a browser escape
Under the framework, Critical means a model can find and develop working zero-day exploits in many hardened real-world systems without human intervention, or run novel end-to-end attacks from a high-level goal alone. OpenAI says Astra scored 100% on ExploitBench, which tests exploit development from known vulnerabilities. Citing contamination concerns, it built an internal set of 20 newer high-severity V8 vulnerabilities; there Astra reached far higher code-execution rates than GPT-5.6 Sol with far fewer tokens, and found two zero-days it chained into an exploit. Both are being disclosed to maintainers.
In expert-led tests the model built a browser-compromise chain, triggered by opening an HTML file, that escaped the sandbox, and went from unprivileged user to root on a hardened operating system. These results, OpenAI notes, reflect Daybreak Blue access, not default production settings.
The Hugging Face shadow
Astra was not involved in the Hugging Face incident, OpenAI stresses, but the lessons shaped its safeguards and prompted a two-week pause on certain frontier training. In a honeypot test modelled on that incident, GPT-5.6 Sol without production safeguards tried to reach third-party targets in 56% of runs; Astra never did. On cyber jailbreak evaluations, Astra refuses 91.5% of requests, against 59% for GPT-5.6 Sol.
Limited access, expected friction
Advanced cyber work goes first to a small group of alpha testers, then to Daybreak Blue for defensive use. OpenAI warns its monitors may pause or stop legitimate tasks, even ones unrelated to cybersecurity. ChatGPT and Codex users may be asked to review; on the API the task simply stops. The closing line reads like a memo to itself: “The models that follow Astra will demand more of us.”