- OpenAI said on Friday it cannot rule out that its upcoming AI model, Astra, has critical cybersecurity capabilities, prompting the startup to pause some internal development and trigger safety protocols.
- Preliminary evaluations over the past several days, along with outside expert assessments, indicated Astra may be capable of performing increasingly sophisticated cyber tasks autonomously, OpenAI said.
- In response to the preliminary findings, OpenAI said it has scaled up security controls and paused internal activities involving Astra that do not meet its newly strengthened security requirements.
- Astra's development will be moved into isolated testing environments with restricted network access and sandboxed execution.
- CEO Sam Altman said on X OpenAI is working to make Astra generally available, as the company does not think it is a good strategy to keep powerful models to a chosen few.
- The model reached its critical cybersecurity threshold, meaning it could independently identify and carry out cyberattacks against well-protected real-world systems, triggering additional safeguards under the Preparedness Framework created in 2023.
- OpenAI clarified that Astra was not involved in the hack targeting the AI platform Hugging Face, and it will partner with government agencies and select AI safety organizations to test the model's capabilities.
- OpenAI is already under scrutiny after a different unreleased model during internal testing the first verifiable incident of an AI lab losing control of its model.
- OpenAI said it is sharing this information because it believes it’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities.
OpenAI has flagged potential critical cybersecurity risks associated with its upcoming AI model, Astra, leading to a pause in development and enhanced security measures.123589
The company announced that preliminary evaluations indicated Astra might autonomously identify and exploit severe software vulnerabilities, known as zero-day exploits, and execute complex cyberattacks against highly secure targets without human intervention.
“While we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out ‘critical’ capability level at this time,” OpenAI stated.
In response to these findings, OpenAI has implemented stricter security controls and moved Astra's development into isolated testing environments with restricted network access.4

CEO Sam Altman emphasized the importance of making powerful models available, stating, “we do not think it is a good strategy to keep powerful models to a chosen few.”
OpenAI clarified that Astra was not involved in a recent hack targeting the AI platform Hugging Face and is partnering with government agencies and select AI safety organizations to rigorously test the model's capabilities.
This disclosure marks a significant moment in the AI sector, where companies often hesitate to announce product delays due to safety and cybersecurity concerns. OpenAI's transparency reflects its commitment to public safety and security amidst growing scrutiny over AI technologies.
“It’s important to be transparent with the public and the safety and security communities about this potential shift in capabilities,” the company added.11
“The model reached the 'critical' threshold under OpenAI's Preparedness Framework, meaning it could autonomously exploit zero-day vulnerabilities. OpenAI has moved Astra to isolated testing environments and will partner with government agencies and select AI safety organizations to test its capabilities.”
