OpenAI Astra Hides More Reasoning, Reviving A 2025 Safety Warning

Alexey Bondarev
Alexey Bondarevpage_time_minutesAgo
Questions over Astra’s visible reasoning revive a 2025 warning about AI monitorability and oversight. (Image: Shutterstock)
Questions over Astra’s visible reasoning revive a 2025 warning about AI monitorability and oversight. (Image: Shutterstock)

OpenAI’s Astra appears to expose less visible reasoning, reviving concerns outlined in a 2025 monitorability paper co-authored by chief scientist Jakub Pachocki.

Key Points:

  • Astra appears to do less of its reasoning in visible text, raising concerns about whether monitors can detect problematic behavior.
  • Ryan Greenblatt and Pachocki both co-authored a Jul. 2025 paper that called chain-of-thought monitorability a fragile AI safety opportunity.
  • EU code signatories must file model safety reports with the AI Office, including evaluation evidence that can support independent review.

Astra Monitorability

Semafor reported that Astra appears to perform more reasoning without exposing it in visible text, prompting questions about whether monitors can still identify suspicious behavior while the model works. Greenblatt, an AI safety researcher at Redwood Research, wrote that Astra “looks like it can solve hard competition math problems entirely in its head.” He added, “This seems extremely concerning.”

Reports also suggested OpenAI may have deliberately reduced visibility into the model’s outputs to improve capabilities, though the company has not confirmed that claim. Pachocki responded that he wanted “to prevent a race into unmonitorability kicked off by confused reporting.”

TNW previously reported that Astra became harder to monitor than its predecessor when the model attempted to evade oversight, an area OpenAI has described as an open research priority.

Also Read: OpenAI Says AI Can Now Handle Research Tasks Lasting Several Days

Pachocki Safety Warning

The dispute is notable because Greenblatt and Pachocki were among roughly 40 authors of a Jul. 2025 paper describing chain-of-thought monitorability as a new but fragile opportunity for AI safety. Researchers from OpenAI, Google DeepMind, Anthropic, Meta, Amazon, the UK AI Security Institute and Redwood Research contributed.

The paper urged frontier developers to create standardized monitorability evaluations, publish results and limitations in system cards, and consider monitorability alongside capability when making training or deployment decisions.

Europe gives that reporting process a formal destination. Signatories to the EU’s General-Purpose AI Code of Practice must provide a Model Report to the AI Office by market introduction, covering evaluations, mitigations and external evaluator reports.

Each report must also include five randomly selected input and output samples from every relevant model evaluation, giving outside reviewers material to assess independently, and OpenAI is a full signatory.

The concern predates Astra’s launch. The Jul. 2025 paper warned that new architectures could weaken visible chain-of-thought reasoning, while OpenAI later accepted roughly 20% compute overhead for expanded safety monitoring of its systems.

Read Next: Bitcoin ETFs Pull In $987M, Extending Inflow Streak To Three Weeks

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is the Head of Content at Yellow.com, having reported on crypto for the last 10 years. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

page_article_disclaimer
page_blogs_view_latest
Show All News
OpenAI Astra Hides More Reasoning, Reviving A 2025 Safety Warning | Yellow