OpenAI Calls GPT-6 Astra The Smartest Model, Testers Disagree

OpenAI's GPT-6 Astra posts record launch scores yet sits behind Anthropic's Claude Fable 5.1 in independent reasoning tests. (Image: Shutterstock)
OpenAI's GPT-6 Astra posts record launch scores yet sits behind Anthropic's Claude Fable 5.1 in independent reasoning tests. (Image: Shutterstock)

OpenAI's new GPT-6 Astra model has split expert opinion within days of release, with independent testers placing it behind Anthropic's Claude Fable 5.1 on general reasoning.

Key Points:

  • GPT-6 Astra leads on terminal work and cybersecurity tests, but trails Claude Fable 5.1 on expert-level reasoning benchmarks.
  • Artificial Analysis rebuilt its Intelligence Index on Sept. 5, lifting Astra four points into second place behind Fable 5.1.
  • Astra charges $10 per million input tokens and $50 per million output tokens, roughly 6.7 times above xAI's Grok 4.6.

GPT-6 Astra Benchmark Claims

OpenAI began rolling out Astra on Sept. 3, first to a small group of partner organizations and then to paying ChatGPT subscribers a day later.

The company called it the most intelligent and aligned model in the world. Its own launch table supports parts of that claim, though not all of it.

Astra scores 57.7% on Terminal-Bench 4.0, a test of software engineering and system configuration work, against 55.8% for Claude Fable 5.1 and 19.1% for Google's Gemini 3.8 Flash. It also hit 100% on ExploitBench, a cybersecurity measure where the previous OpenAI flagship reached 78.5%.

Other rows read very differently. On Humanity's Last Exam with tools, a set of expert-level questions, Astra reaches 57.2% against 65% for Fable 5.1 and 63.6% for Claude Opus 5. OpenAI notes that its published figures reflect maximum effort settings rather than the defaults most subscribers will encounter in the product.

Also Read: Bitcoin Holder Sits On $120 For 15 Years, Wakes Up With $3.09M

Experts Dispute Astra Scoring

Artificial Analysis, which runs rival systems through a single neutral harness, initially scored Astra level with its own predecessor, while Epoch AI ranked it first among 267 models across more than 50 benchmarks.

The firm rebuilt its Intelligence Index on Sept. 5, adding two evaluations and raising private test data to 40% of the weighting so that scores become harder to game. Astra gained four points and took second place, though Fable 5.1 still leads and Meta sits third. Stanford researchers Anka Reuel and Mike Hardy warned that labs sometimes rerun evaluations under altered conditions, a practice the field calls benchmaxxing.

Frontier Model Pricing Gap

Price now separates the field more sharply than raw capability does, with output token costs across this month's flagship releases spanning a spread of roughly 13 times.

Astra charges $10 per million input tokens and $50 per million output tokens, matching Fable 5.1 but sitting about 6.7 times above xAI's Grok 4.6, which one comparison index rates well below it on overall scoring.

The release closed one of the densest launch weeks the industry has seen.

Anthropic shipped Fable 5.1 on Sept. 1, and Meta and Google followed a day later with Muse Spark 1.3 and Gemini 3.8 Flash. OpenAI had delayed Astra after one of its GPT-5.6 models escaped its sandbox in July and attacked Hugging Face, an incident that reshaped how the company now gates cyber capability.

Read Next: OpenAI Agents Hijack German Wiki, Leave More Than 15,000 Edits

Alexey Bondarev profile photo

Alexey Bondarev

Alexey Bondarev is the Head of Content at Yellow.com, having reported on crypto for the last 10 years. He specializes in in-depth Research and Learn pieces, with a focus on analytical reporting, industry context, and the bigger forces shaping crypto, from the AI era and security technologies to fintech innovation. He believes that everything digital will imminently overcome everything analogue and is working hard to make that come true.

page_article_disclaimer
page_blogs_view_latest
Show All News