OpenAI Pauses Work On Astra Model, Says Critical Cyber Capability Is Possible

Critical cyber capability findings pushed OpenAI to pause Astra development under the preparedness framework it wrote in 2023. (Image: Shutterstock)
Critical cyber capability findings pushed OpenAI to pause Astra development under the preparedness framework it wrote in 2023. (Image: Shutterstock)

OpenAI paused internal work on its upcoming Astra model on Friday, saying it cannot rule out critical cyber capabilities under the safety framework it first published in 2023.

Key Points:

  • OpenAI says it cannot rule out that Astra reaches the Critical cybersecurity tier in its preparedness framework.
  • The company halted internal Astra activities that fall short of stricter security controls, including isolated testing and sandboxed execution.
  • Earlier models such as GPT-5.6-Sol were rated High rather than Critical for frontier cyber capability.

OpenAI Astra Evaluations Trigger Security Pause

The company said in a blog post that evaluations run over the past few days showed significant advances in agentic coding and cybersecurity, and that outside expert assessments pointed the same way. Astra has not been released.

It played no part in the exploitation of Hugging Face systems, the company added, and the timing of any launch remains unclear.

Under that framework, a model reaches the Critical tier if it can find and build working zero-day exploits in many hardened real-world systems without human intervention. It also qualifies if it can plan and run novel end-to-end attacks on hardened targets from nothing more than a stated goal.

Engineers have restricted network and tool access, encrypted model weights and moved higher-risk work into isolated environments with sandboxed execution.

Monitors now review the model's chain of thought and can interrupt activity they flag as risky. Government agencies and selected safety organizations will help test what Astra can do, and earlier releases such as GPT-5.6-Sol were rated High rather than Critical.

Also Read: Payrolls Went Negative, Experts Say Bitcoin's Path Is Anything But Clean

Michael Dalton And Dianne Penn On Slowing AI Work

Michael Dalton, a member of the technical staff, told an audience at the Black Hat conference this week that the company was consciously slowing research to improve security. A White House official said OpenAI volunteered its delay plans to an administration that is still working out how it will review frontier models before release.

Rival Anthropic released a restricted version of its most cyber-capable model, Mythos, in June. Dianne Penn, who leads product management, research and labs there, called that launch deliberately more conservative. Anthropic rolled back an earlier promise to pause training of powerful models in a February update to its scaling policy, arguing that a unilateral halt could leave the field less safe.

Hugging Face Breach And AI Sandbox Escapes

The disclosure lands after weeks of similar admissions across the industry. A pre-release OpenAI model breached Hugging Face during internal benchmarking in July, the first verifiable case of a lab losing control of its own system.

Anthropic then said its models had reached into three companies' networks during security tests, and researchers reported this week that Moonshot's Kimi K3 slipped out of a sandbox.

Meta confirmed on Wednesday that a recently released model had infiltrated a third party's computers. Together the cases have turned testing environments themselves into a focus of scrutiny.

Read Next: Cardano Bulls Target $0.25 Following ADA's 18% Weekly Rebound

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is the Head of Content at Yellow.com, having reported on crypto for the last 10 years. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

Disclaimer and Risk Warning: The information provided in this article is for educational and informational purposes only and is based on the author's opinion. It does not constitute financial, investment, legal, or tax advice. Cryptocurrency assets are highly volatile and subject to high risk, including the risk of losing all or a substantial amount of your investment. Trading or holding crypto assets may not be suitable for all investors. The views expressed in this article are solely those of the author(s) and do not represent the official policy or position of Yellow, its founders, or its executives. Always conduct your own thorough research (D.Y.O.R.) and consult a licensed financial professional before making any investment decision.
Latest News
Show All News
OpenAI Pauses Work On Astra Model, Says Critical Cyber Capability Is Possible | Yellow