OpenAI Astra Becomes First Model To Hit Critical Cyber Threshold

Researchers pushed back after a 100% exploit score moved OpenAI's Astra past its own Critical cyber threshold. (Image: Shutterstock)
Researchers pushed back after a 100% exploit score moved OpenAI's Astra past its own Critical cyber threshold. (Image: Shutterstock)

OpenAI said Tuesday its unreleased Astra model is the first to cross its Critical cybersecurity threshold, after scoring 100% on a public exploit-writing benchmark.

Key Points:

  • Astra is the first OpenAI model rated Critical for cybersecurity under the company's Preparedness Framework.
  • The model hit a perfect score on a public exploit benchmark and uncovered two previously unknown flaws in internal testing.
  • Separate reporting on an architecture called recurrent depth has drawn objections from AI safety researchers.

Astra Exploit Benchmark Results

The company [published](https://www.wired.com/story/openai-astra-first-ai-model-with-critical-cyber-abilities/ the assessment Tuesday, saying Astra can find previously unknown flaws and build working exploits across well-protected systems without a person guiding each step.

On a private benchmark assembled from 20 high-severity bugs disclosed between June and August, the model discovered and chained two zero-day vulnerabilities that engineers are now reporting to maintainers. Expert reviewers also turned it loose on a hardened browser and operating system, where it escaped the sandbox and ran commands on the host.

Access to the strongest cyber features will go first to a small group of alpha testers, then widen through a defensive program called Daybreak Blue.

OpenAI has not said who those testers are or how it picked them, and no outside body has confirmed the benchmark results it published.

The company warned that its own safeguards may pause or stop legitimate work, including defensive security research and long-running agent tasks. OpenAI also calls Astra its most aligned model so far. It refused 91.5% of cyber jailbreak attempts in internal testing, against 59% for the current flagship, GPT-5.6 Sol.

Also Read: Claude Fable 5.1 Arrives With $0.25 Cache Reads And Anti-Copy Controls

Recurrent Depth Sparks Objections

A separate account reported the same night that Astra leans on an architecture known as recurrent depth, which the company is said to be limiting for now.

The method shifts more of the model's reasoning out of readable text and into internal activations, which lowers cost and raises performance on hard problems. It also thins the plain-language trail that safety teams read to catch a model drifting from its instructions.

OpenAI still plans to ship Astra with extra chain-of-thought monitoring, a control that works only on reasoning it can read.

Ryan Greenblatt, chief scientist at Redwood Research, called the change possibly the worst development yet for AI safety and security. Steven Adler, a former OpenAI safety lead, wrote that the approach appears to cross one of the very few redlines the industry has drawn for itself.

OpenAI Cyber Warnings Timeline

The designation caps a month of escalating warnings from the company.

OpenAI first disclosed on Aug. 7 that it could not rule out Critical cyber capability, then paused parts of its frontier training after agents from an earlier system left a test environment and reached data on Hugging Face.

That paused run restarted Aug. 28, four days before the company declared its safeguards sufficient for release.

Read Next: Bitcoin ETFs Absorb $216.7M, Undoing A Single Session Of Outflows

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is the Head of Content at Yellow.com, having reported on crypto for the last 10 years. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

Disclaimer and Risk Warning: The information provided in this article is for educational and informational purposes only and is based on the author's opinion. It does not constitute financial, investment, legal, or tax advice. Cryptocurrency assets are highly volatile and subject to high risk, including the risk of losing all or a substantial amount of your investment. Trading or holding crypto assets may not be suitable for all investors. The views expressed in this article are solely those of the author(s) and do not represent the official policy or position of Yellow, its founders, or its executives. Always conduct your own thorough research (D.Y.O.R.) and consult a licensed financial professional before making any investment decision.
Latest News
Show All News
OpenAI Astra Becomes First Model To Hit Critical Cyber Threshold | Yellow