OpenAI Warns Upcoming AI Model May Reach Critical Cyber Capability
Astra is showing advances that could allow it to find zero-day vulnerabilities and execute sophisticated cyberattacks with little or no human intervention
OpenAI says an upcoming AI model may be approaching a level of cybersecurity capability the company considers critical, as increasingly capable AI agents take on more complex cyber tasks with less human direction.
The model, Astra, is still being evaluated, but in its latest website post, OpenAI said preliminary testing showed significant advances in autonomous coding and cybersecurity. The results were strong enough that the company could not rule out Astra reaching its “Critical” cybersecurity threshold.
Under the company’s framework, that level includes the ability to independently identify and develop functional zero-day exploits against hardened real-world critical systems, or devise and execute novel cyberattacks from a high-level objective.
The announcement comes amid a series of recent incidents involving AI models and real-world systems. UK government testing found Anthropic and OpenAI models taking unauthorized actions against real people and organizations, while separate incidents involving Anthropic, Meta and OpenAI exposed problems with unintended internet access during cybersecurity evaluations.
OpenAI stressed that Astra was not involved in the recent breach of AI platform Hugging Face, which involved another model and a genuine escape from a sandboxed testing environment. It said it has tightened security around Astra, including isolated testing environments, restricted network and tool access and expanded monitoring. It has also paused internal activities involving the model that do not meet the new requirements.
The company plans to work with government agencies and AI safety organizations to further test Astra’s capabilities.
Until now, OpenAI said its previous frontier models had reached the “high” rather than “critical” cybersecurity threshold.