- OpenAI said Astra has reached the critical cybersecurity-capability threshold under its preparedness framework and plans to make the model available in the near future.
- OpenAI classified Astra as its first model at that level, meaning the company says it can independently find unknown vulnerabilities and build exploit chains against hardened systems.
- During evaluations, Astra discovered and exploited two zero-day vulnerabilities in the V8 engine as part of an exploit-chain attack, according to OpenAI.
- Astra refused 91.5% of dangerous requests in tests, compared with 59% for GPT-5.6 Sol, the company said.
OpenAI said Astra has reached the threshold for critical cybersecurity capabilities under its preparedness framework and plans to make the model available in the near future. The designation matters because OpenAI says Astra is its first model capable of finding previously unknown vulnerabilities and developing ways to exploit them in well-protected systems without step-by-step human guidance.
OpenAI said it delayed some stages of Astra’s development and release over the past few weeks to strengthen and test safeguards against cyber misuse and unauthorized actions.
Tests assess zero-day and attack capabilities
Under OpenAI’s preparedness framework, a model reaches the critical threshold if it can find and create functional zero-day exploits for many hardened critical systems without human intervention. A model can also meet the threshold by developing and executing end-to-end novel cyberattack strategies against protected targets from only a high-level objective.
OpenAI evaluated Astra through automated public and private benchmarks and expert-led experiments. On ExploitBench, the model recorded a 100% score in an assessment of its ability to develop exploits for known vulnerabilities.
The company also created an internal benchmark called ExploitBench — Internal Port, which covers 20 recently disclosed high-severity vulnerabilities in the V8 engine. According to OpenAI, Astra achieved a significantly higher level of arbitrary code execution than GPT-5.6 Sol while using far fewer output tokens.
During that evaluation, Astra also discovered and exploited two zero-day vulnerabilities as part of an exploit-chain attack, OpenAI said. The company said it is disclosing details to the developers of the affected software.
In separate experiments involving a hardened browser and operating system, Astra found previously unknown vulnerabilities and combined them into working exploit chains, according to OpenAI.
OpenAI strengthens safeguards
OpenAI identified two principal risks associated with models at this capability level. The first is that malicious actors could use Astra to create exploits for unknown flaws in protected critical systems or conduct complex attacks. The second is that alignment problems could cause the model itself to take unauthorized actions.
The company said it further strengthened Astra’s model hardening within its safeguards system and improved its handling of cross-session context. In tests, Astra refused 91.5% of dangerous requests, compared with 59% for GPT-5.6 Sol.
OpenAI said it applies more conservative behavioral limits and enhanced monitoring for potential cyber abuse to higher-risk accounts. It is also continuing internal and external red-teaming, regression testing and efforts to identify new jailbreaks.
The company plans to operate a 24-hour rapid-response program to investigate and address newly identified vulnerabilities.
Initial access limited to testers
Astra’s most advanced cybersecurity capabilities will initially be available only to a limited group of testers. OpenAI said it will first give a small group of alpha testers access for cutting-edge cybersecurity work before expanding defensive-use access through the Daybreak Blue program.
The company warned that the stronger safeguards may sometimes misclassify legitimate activity as potential cyber abuse or unauthorized behavior. Such classifications could slow, pause or stop a task.
OpenAI and another 155 companies previously called for stronger AI cybersecurity.
Source: Incrypted
