OpenAI Pauses Some Astra Work Over Cyber Risks
Preliminary tests suggest Astra may meet OpenAI’s Critical cybersecurity threshold, prompting tighter controls on testing, networks and model access.
Topics
News
- Temasek, Seraphim Back Pixxel in $100 Million Round
- UN Rights Chief Warns Advanced AI Could Threaten Humanity
- Tata-owned JLR to Cut About 4,000 Bobs Amid Tariff, China Pressure
- Anthropic Walks Away From $6 Billion Decart Deal
- Hyundai India Targets 20% Women in Executive Workforce by 2030
- Security Breach Drains $320 Million in Bitcoin From Liquid Network
OpenAI has paused some internal work involving its forthcoming Astra model after preliminary evaluations showed performance strong enough that the company said it could not rule out its highest cybersecurity capability threshold.
OpenAI said Friday that recent tests showed “significant advancements in agentic coding and cybersecurity,” prompting it to strengthen safeguards while Astra remains under development.
Under OpenAI’s Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can develop functional zero-day exploits against many hardened real-world critical systems without human intervention, or independently devise and carry out novel attacks against hardened targets from a high-level goal.
“Our preliminary evaluations indicate strong enough performance that we cannot rule out Critical capability level at this time,” OpenAI said.
The company has paused internal Astra activities that do not meet strengthened security requirements. The measures include isolated testing environments, restricted network and tool access, stronger protection and encryption of model weights, sandboxed execution and additional monitoring.
OpenAI said it has also introduced monitoring for risky actions across Astra’s agentic applications and will work with government agencies and selected AI safety organizations to further test the model.
Astra has not been released, and OpenAI stressed that it was not involved in the July security incident involving Hugging Face.
That incident involved GPT-5.6 Sol and an internal research prototype being tested with reduced cyber safeguards. The models exploited a previously unknown vulnerability to gain internet access from OpenAI’s test environment and ultimately compromised Hugging Face’s production infrastructure. OpenAI later said the prototype was not intended for public release.
The Astra disclosure comes amid a series of incidents involving advanced AI agents during cybersecurity testing.
Meta said last week that one of its models exploited a vulnerability in a third-party service after a testing configuration inadvertently gave it internet access. Anthropic has also disclosed instances in which its models accessed external organizations during cyber evaluations.
Britain’s AI Security Institute reported on August 4 that agents powered by OpenAI and Anthropic models took unauthorized actions involving real people and organizations during a cybersecurity exercise. Researchers had deliberately provided internet access and disabled some safeguards to test the models’ maximum capabilities.
The institute recorded 19 out-of-scope actions across 10 of 122 test runs. These included an attempted malicious code contribution to an open-source project and efforts to influence human developers. AISI said it found no evidence of resulting real-world harm.
Unlike those incidents, OpenAI’s latest announcement concerns what Astra demonstrated in capability testing rather than a disclosed breach involving the model.
OpenAI said it will continue benchmarking Astra as it assesses the model’s capabilities and the safeguards needed before deployment.


