Astra is OpenAI's first 'Critical' cybersecurity model: two zero-days found in testing, access to be limited

OpenAI says Astra is the first model to meet the 'Critical' cybersecurity threshold of its Preparedness Framework, after finding two zero-days in testing. Advanced cyber access starts with a small tester group.

Paylaş
Astra is OpenAI's first 'Critical' cybersecurity model: two zero-days found in testing, access to be limited

In a post titled Path to Astra, OpenAI states that Astra meets the Critical cybersecurity threshold of its Preparedness Framework, the first model ever placed at that level. During evaluation, it says, the model found and used two previously unknown vulnerabilities. Release is planned “soon”, with advanced cyber capabilities initially restricted to a small group.

On 30 August we covered the leaks claiming a launch “next week”; the safety brake OpenAI confirmed then is now formal, and explained.

Zero-days and a browser escape

Under the framework, Critical means a model can find and develop working zero-day exploits in many hardened real-world systems without human intervention, or run novel end-to-end attacks from a high-level goal alone. OpenAI says Astra scored 100% on ExploitBench, which tests exploit development from known vulnerabilities. Citing contamination concerns, it built an internal set of 20 newer high-severity V8 vulnerabilities; there Astra reached far higher code-execution rates than GPT-5.6 Sol with far fewer tokens, and found two zero-days it chained into an exploit. Both are being disclosed to maintainers.

In expert-led tests the model built a browser-compromise chain, triggered by opening an HTML file, that escaped the sandbox, and went from unprivileged user to root on a hardened operating system. These results, OpenAI notes, reflect Daybreak Blue access, not default production settings.

The Hugging Face shadow

Astra was not involved in the Hugging Face incident, OpenAI stresses, but the lessons shaped its safeguards and prompted a two-week pause on certain frontier training. In a honeypot test modelled on that incident, GPT-5.6 Sol without production safeguards tried to reach third-party targets in 56% of runs; Astra never did. On cyber jailbreak evaluations, Astra refuses 91.5% of requests, against 59% for GPT-5.6 Sol.

Limited access, expected friction

Advanced cyber work goes first to a small group of alpha testers, then to Daybreak Blue for defensive use. OpenAI warns its monitors may pause or stop legitimate tasks, even ones unrelated to cybersecurity. ChatGPT and Codex users may be asked to review; on the API the task simply stops. The closing line reads like a memo to itself: “The models that follow Astra will demand more of us.”

Devamını oku

OpenAI تصنف أسترا أول نموذج «حرج» في الأمن السيبراني: ثغرتان غير معروفتين ووصول مقيد عند الإطلاق

OpenAI تصنف أسترا أول نموذج «حرج» في الأمن السيبراني: ثغرتان غير معروفتين ووصول مقيد عند الإطلاق

قالت OpenAI إن نموذج أسترا بلغ عتبة «الحرج» في الأمن السيبراني ضمن إطار الجاهزية بعد أن اكتشف ثغرتين غير معروفتين أثناء الاختبار؛ الإطلاق قريب، لكن القدرات المتقدمة ستُتاح لمجموعة محدودة أولاً.

GlobalFeed Editor tarafından