OpenAI's new AI can find and exploit security vulnerabilities on its own

On September 1, OpenAI introduced more information about Astra, an AI model being developed and preparing for release, and said it had to strengthen protective measures after evaluating the model's cybersecurity capabilities.
According to the company, Astra is OpenAI's first model to exceed the "Critical" threshold for cybersecurity capabilities, according to the standard framework the company uses to evaluate AI capabilities that pose high risks.
At this level, OpenAI estimates that when given the right tools and access, the model can "find previously unknown security vulnerabilities and develop ways to exploit them on multiple protected systems, without needing step-by-step guidance from humans."
"This is the first model we've specified at this level and requires stronger protections during development and before release," OpenAI introduced.
In some of the tests mentioned by OpenAI in the introduction, the company said Astra achieved 100% on ExploitBench - an assessment of the ability to develop exploit tools from available vulnerabilities. For example, Astra exploits these vulnerabilities significantly better than the GPT-5.6 Sol model, while requiring less output data to complete the task. During testing, Astra itself discovered and exploited two previously unknown zero-day vulnerabilities, combining them into an attack chain. OpenAI said it is coordinating with software maintainers to announce and handle these two vulnerabilities.

In another review, Astra built an exploit chain that starts in the browser, bypasses the "sandbox" - a quarantine area that limits code running in the browser from reaching external systems - and then executes commands on the server. In another test, the model found vulnerabilities on a protected operating system and combined them to elevate privileges from a regular user account to root, which is the highest level of control on the system.
The announcement of the Astra trials comes after OpenAI was met with cybersecurity concerns. The investigation published in late August found that about 700 actors coordinated to attack the platform. However, Astra was not involved in the incident and OpenAI confirmed that the models planned for release were also not involved in the attack. The company later added monitoring measures, increasing its ability to detect unauthorized actions and training the model to reject requests that potentially support cyberattacks. In tests published by the company, Astra rejected 91.5% of unauthorized network attack requests, compared to 59% for GPT-5.6 Sol.

OpenAI said it will release Astra, but not widely, but will make it limited to a select group of partners and testers. The goal is to use these capabilities primarily for defensive purposes, helping to detect and address weaknesses that humans might otherwise miss.
The ability for AI to find and exploit vulnerabilities on its own is also becoming a source of competition among model development companies. Previously, Anthropic was also developing the Claude Mythos line for cybersecurity tasks. In April, the company said the Mythos Preview version had discovered thousands of serious vulnerabilities in many operating systems and browsers. On September 1, Anthropic continued to introduce Claude Mythos 5.1, a new version of this model line, with improvements in cybersecurity and biological research capabilities. Access currently remains limited to a select group of organizations.
Luu Quy