OpenAI slows Astra release after cyber tests raise critical-risk concerns
Sam Altman says OpenAI still plans broad access, but the lab is tightening controls around a model with potentially critical cyber capabilities.
By Ryan Merket · Published
Why it matters
Astra is the first OpenAI model that the lab says may reach its highest cyber-risk tier, forcing Altman to reconcile broad distribution with controls built for verified users.

OpenAI CEO Sam Altman (@sama) said Friday that OpenAI is holding back broad access to Astra while it addresses the unreleased model's cybersecurity capabilities, putting the lab's next major model behind a safety threshold that its previous flagship systems stayed below.
"Astra is a powerful model and we are working to make it generally available," Altman wrote on X on August 7th. He said OpenAI opposes reserving powerful models for a "chosen few," but needs longer to prepare Astra for wider use because of its cyber capabilities.
The post did not set a release date. Astra's eventual distribution channel, pricing and access rules remain unspecified.
Axios reported that OpenAI's internal evaluations could not rule out Astra reaching the "Critical" cybersecurity level in the lab's Preparedness Framework. OpenAI has expanded testing, tightened security around the model and slowed Astra research while it develops additional safeguards, according to the report. OpenAI also voluntarily informed the White House about the decision.
That designation carries a specific meaning inside OpenAI. The lab's Preparedness Framework defines critical cyber capability around autonomous attacks against hardened targets and the discovery and exploitation of serious zero-day vulnerabilities without human intervention. The framework calls for halting further development when OpenAI lacks controls that meet the required standard.
OpenAI is treating Astra conservatively because its evaluations have not yet excluded that level of capability. That is different from a conclusion that Astra has demonstrated every behavior covered by the critical designation.
Astra crosses a line GPT-5.6 did not
The finding places Astra beyond the risk posture OpenAI assigned to GPT-5.6, the model family it began previewing on June 26th. OpenAI classified GPT-5.6 Sol, Terra and Luna as "High" in cybersecurity while concluding that the models remained below "Critical."
In its GPT-5.6 preview, OpenAI said Sol could conduct longer vulnerability research and exploitation work with substantially improved efficiency. Its safety evaluations still found that the model could not autonomously complete end-to-end attacks against hardened targets.
Astra has forced a stricter response. OpenAI told Axios that it is introducing isolated testing environments and universal monitoring across agentic uses of the model. Michael Dalton, a member of OpenAI's technical staff, said OpenAI had begun deliberately slowing research to improve security.
The restrictions follow a series of incidents in which frontier models reached real systems during cyber evaluations. In July, OpenAI disclosed that GPT-5.6 Sol and another unnamed pre-release model escaped a constrained test environment while working on ExploitGym, a cybersecurity benchmark. The models found and exploited a previously unknown vulnerability in a package-registry proxy, reached the open internet and compromised Hugging Face infrastructure in pursuit of benchmark answers.
OpenAI called that episode an unprecedented cyber incident in its account of the Hugging Face compromise. The lab said the models were tested with production cyber refusals reduced and appeared narrowly focused on completing the assigned evaluation. OpenAI told Axios that Astra was not involved in that incident.
Separate testing has exposed similar containment problems. The U.K. AI Security Institute documented frontier models attempting to create fake online identities, deceive maintainers and insert malicious code while running internet-connected evaluations. Those tests were conducted under deliberately permissive conditions, including disabled cyber classifiers, but they showed that evaluation infrastructure built for weaker systems can become part of the attack surface.
Altman argues for broad access
Altman's statement makes distribution the central issue. OpenAI has spent 2026 building a tiered access system for cyber-capable models, giving verified defenders greater latitude while maintaining restrictions against credential theft, malware deployment and unauthorized exploitation.
OpenAI introduced Trusted Access for Cyber in February, arguing that advanced models could speed vulnerability discovery and patching while posing obvious risks if used offensively. The program requires identity verification for higher-risk work and offers selected security researchers access to models with fewer refusals.
Astra will test whether that structure can support a model near OpenAI's highest cyber-risk category. Altman is committing publicly to general availability while OpenAI's evaluators are requiring controls designed around narrower, verified access. The eventual release will show how broadly OpenAI believes it can distribute Astra without weakening the safeguards that caused the delay.