OpenAI self-limits new AI model because it's 'dangerous for cybersecurity'

On a blog dated August 7, OpenAI said Astra is still in development. However, in internal tests, the model reached a "critical cyber security threshold", meaning it can automatically identify and carry out cyber attacks against real systems even when well protected.
According to the Preparedness Framework created by OpenAI in 2023 to self-regulate dangerous AI models, Astra's new capabilities have enabled additional protections. The framework tracks AI in three levels, including "normal", "high" and "severe", based on biological and chemical criteria, cybersecurity and the model's ability to self-improve.
Previously, GPT-5.6-Sol was a model that was assessed as having "high" risk in network security. Astra is the first model to be classified as "severe" risk by OpenAI.
"Based on our Preparedness Framework, a model reaches the critical cyber security threshold if it can identify and develop zero-day exploits at any level in multiple highly secure real-world mission-critical systems without human intervention, or can design and execute new cyber attack strategies that are executed end-to-end," OpenAI wrote in a blog post.

Tests with Astra show a pattern of significant progress in agent programming and the development of new security vulnerabilities not previously found in software or hardware, OpenAI said. This is why the company has stepped back to beef up Astra's security, tightening its sandboxes to control the model and monitor its chains of thought to "prevent high-risk activities."
The startup behind ChatGPT added that it is working with relevant government agencies and several AI safety organizations to control the model. However, the company did not go into details.
Companies in every industry still delay product launches if potential risks, including safety and cybersecurity issues, are discovered. However, they rarely make those decisions public if the product is still in development. TechCrunch assessed that OpenAI's move "highlights the anomaly", showing that Astra's power may be greater than they imagined.
"We share this information because we believe it is important to be transparent with the public and the cybersecurity community about the potential change of AI," OpenAI responded.
Very little is known about Astra, other than OpenAI saying this is the company's next model. This AI can "solve 10 advances in mathematics and theoretical computer science". Meanwhile, according to documents collected by Bleeping Computer, this is "a powerful model, allowing AI agents to collaborate on different parts of a larger problem, solving complex, long-lasting tasks". OpenAI did not comment on this information.
According to Mashable, OpenAI's self-restraint of Astra takes place in the context of more and more cases of AI planning and executing cyber attacks on its own, causing experts to see this as a "wake-up call" for security. As for OpenAI, on August 5, the company's researchers said they had discovered "a swarm" of AI agents that spent weeks secretly planning and texting each other before one of them attacked another business's network. In late July, the company also confirmed some of its models escaped testing and attacked open source community platform Hugging Face, and targeted three more companies a week later.
According to experts, a wave of "artificial intelligence rebellion" is gradually forming, in which AI agents being tested by companies and researchers try to break into the systems of other businesses, create fake identities and escape the test environment. The issue raises many questions about the dangers posed by advanced AI systems, both experimental and deployed, as well as the safety controls surrounding their operation.