OpenAI has announced ahead of the impending release of Astra, its latest frontier model, that it is the first model to reach the Critical cybersecurity threshold under its Preparedness Framework.
Accordingly, the company has said the model will be released with tighter controls around its most advanced cyber capabilities.
The development points to a growing challenge for frontier AI developers: as models become capable of carrying out more sophisticated cybersecurity tasks with less human direction, the criteria for deploying them safely are also having to change.
Astra Pushes AI Cybersecurity Further
OpenAI’s evaluations show Astra making substantial gains over GPT-5.6 Sol across several cybersecurity tasks. The model scored 100% on ExploitBench and achieved higher arbitrary-code-execution rates than its predecessor on a newer internal benchmark, while using significantly fewer output tokens.
The internal benchmark was built around 20 high-severity V8 vulnerabilities and was designed to test whether Astra could move beyond identifying known weaknesses and independently develop the exploit chains needed to achieve arbitrary code execution. OpenAI says Astra demonstrated significantly greater capability and token efficiency than GPT-5.6 Sol in the evaluation.
The results included two particularly significant discoveries. During evaluation, Astra identified and used two previously unknown vulnerabilities as part of an exploit chain. OpenAI said it is disclosing both vulnerabilities to the affected maintainers.
The company also observed Astra building a browser compromise chain that escaped the sandbox and executed commands on the host machine. The findings demonstrate the extent to which Astra can carry out sophisticated cybersecurity tasks with less step-by-step human guidance.
OpenAI’s classification is the first time one of its models has reached the Critical cybersecurity threshold under its Preparedness Framework. The designation applies when a model can identify and develop functional zero-day exploits across a range of hardened real-world systems without human intervention, or plan and execute novel end-to-end attacks against hardened targets from a high-level objective.
Why Astra Is Raising the Bar
What separates Astra from earlier AI security tools is not simply its ability to find vulnerabilities. It is the extent to which it can connect the stages that follow, moving from discovering a weakness to developing and using an exploit with considerably less human direction.
That distinction changes the role AI can play in cybersecurity. Earlier systems could assist researchers by analyzing code, identifying potential weaknesses or generating exploit ideas, but the human operator remained responsible for deciding how those pieces fitted together and what to do next. Astra’s evaluations suggest that more of that process can now be handled by the model itself.




