GPT-6 Astra launch and benchmark manipulation controversy
TECH

GPT-6 Astra launch and benchmark manipulation controversy

30+
Signals

Strategic Overview

  • 01.
    OpenAI released GPT-6 Astra in a limited preview on September 3, 2026, rolling it out to ChatGPT Plus, Pro, Business, and Enterprise users plus the API, Azure, and AWS Bedrock, with a launch blog post touting benchmark gains that would quickly draw scrutiny.
  • 02.
    Within a day of launch, reporting revealed OpenAI had quietly revised at least five published evaluation figures - including hallucination rate, an ExploitBench cybersecurity score, and ARC-AGI-3 - in ways that favored Astra, with some numbers continuing to shift after publication.
  • 03.
    The most consequential single figure is ARC-AGI-3: OpenAI reported 99.9% using its own Provider Adapter harness, versus 62.7% when the benchmark's creator, ARC Prize, ran Astra through its own neutral, model-agnostic harness.
  • 04.
    Independent benchmarking firm Artificial Analysis found Astra's overall Intelligence Index essentially tied with its predecessor and behind Anthropic's Claude Fable 5.1, contradicting the scale of the gains implied by OpenAI's own published charts.

Deep Analysis

The Harness Gap: How a 62.7% Score Became 99.9% - and an AGI Claim ARC Prize Rejected

The Harness Gap: How a 62.7% Score Became 99.9% - and an AGI Claim ARC Prize Rejected
OpenAI's Provider Adapter harness produced a far higher ARC-AGI-3 score than the benchmark creator's own neutral test.

OpenAI's headline claim that GPT-6 Astra hit 99.9% on the ARC-AGI-3 benchmark looks far less clean once measured against ARC Prize's own numbers. Under OpenAI's proprietary Provider Adapter - a wrapper that preserves the model's opaque reasoning state across requests and compacts long conversations - Astra scored 99.9% at a cost of $18,817. Under ARC Prize's standard, model-agnostic harness at maximum reasoning, the same model scored 62.7% for $26,098[1]. That is not a rounding difference; it is two different tests wearing the same benchmark's name. OpenAI's own leadership leaned into AGI language around the launch - President Greg Brockman said 'I think it's not unreasonable to feel that we are now in the AGI era'[4]- which makes what came next pointed rather than incidental: ARC Prize itself, the benchmark's creator, explicitly declined to endorse that framing, with co-founder Mike Knoop writing that 'we lack evidence to call this AGI yet'[1]. Co-creator Francois Chollet took a more optimistic read, arguing Astra shows genuine internal symbolic-modeling capability that previously required an external harness to unlock, and moved up his own AGI timeline forecast as a result[2]. Both views can be true at once: the model may be genuinely stronger, but the reported number is a harness artifact, not an apples-to-apples capability score, and it remains well short of the AGI framing OpenAI's own president floated.

A Blog Post That Changed Its Mind Four Times

Internet Archive snapshots of OpenAI's own launch page show the published numbers moving during and after release, not just once. Astra's headline hallucination rate read 4.2% in the first snapshot at 2:23 p.m. ET, was revised down to 2% in a later version, then reverted to 4.2% - with the GPT-5.6 Sol comparison figure also shifting before settling back at 12.2%[3]. GPT-5.6 Sol's ExploitBench cybersecurity score doubled mid-launch, from 5.5% to 11.5%; OpenAI later said it was investigating reverting the number because it reflects a reasoning tier not commercially available[3]. The ARC-AGI-3 score sent to press under embargo was 98.6%, then published live at 99.99%[3], and even a competitor's number moved: Anthropic's Fable 5.1 FrontierMath comparison score dropped nearly 10 points, from 87.8% to 78%, between snapshots[3]. Snorkel AI's Vincent Sunn Chen offered a more charitable read, characterizing the edits as typical of final launch logistics rather than manipulation[3], while Stanford researchers Anka Reuel and Mike Hardy argued the system card lacked adequate documentation for how internal evaluations like hallucination rate are produced, and suggested the edits conveniently favored the marketing narrative[3].

OpenAI's 'Critical' Cybersecurity Claim Rests on the Same Shaky Number

OpenAI's launch blog post claims Astra is the company's first model to reach the 'Critical' level of cybersecurity capability under its own Preparedness Framework[4]- the internal framework OpenAI uses to gate additional safety review before a model ships. That classification is not a marketing footnote; it is meant to be the trigger for extra scrutiny on a model's real-world risk. Yet the primary evidence behind Astra's cybersecurity story, the ExploitBench score, is exactly the kind of number this launch has already shown to be unstable: GPT-5.6 Sol's ExploitBench figure doubled between versions of OpenAI's own blog post, from 5.5% to 11.5%, with OpenAI attributing the jump to a reasoning tier not commercially available[3]. When the underlying benchmark feeding a safety-relevant capability threshold moves that much within the same launch cycle, the 'Critical' designation built on top of it inherits the same credibility problem as every other disputed figure in this launch: it is difficult to know whether the classification would have held at the benchmark's earlier, lower number, or whether it was set against the version that happened to clear the bar.

What Independent Benchmarks Actually Show

What Independent Benchmarks Actually Show
Independent testing places Astra roughly level with its predecessor, at odds with OpenAI's own published charts.

Set against third-party testing, Astra's story looks far less dramatic than OpenAI's own charts suggest. Artificial Analysis's Intelligence Index puts Astra at 61.2, essentially tied with its predecessor GPT-5.6 Sol at 60.9, and behind Anthropic's Claude Fable 5.1 at 65.7[5]- a picture hard to reconcile with a launch pitched as a generational leap. On hallucination specifically, Artificial Analysis's own testing found Astra hallucinating in 51% of cases at maximum effort, down from 92% for Sol - a real improvement, but nowhere near the sub-5% figures OpenAI published[6]. Independent measurement and OpenAI's own published charts are telling two different stories about the same launch, and only one of them is reproducible outside OpenAI's own infrastructure.

This Is the Third Time, Not the First

GPT-6 Astra's benchmark dispute is not an isolated incident for OpenAI. In late 2024, its o3 model's claimed 25% score on the FrontierMath benchmark could not be reproduced independently, with Epoch AI's own testing landing closer to 10%[7]. It later emerged that OpenAI had secretly funded and had advance access to the FrontierMath dataset before that launch, without disclosing the arrangement to the academic contributors involved[8]. More recently, GPT-5.5 drew criticism when users discovered its premium Extended Thinking mode was being silently downgraded for many queries[9]. Reaction to Astra split hard along platform lines once the pattern became visible. On X, an early hype-driven tweet touting the raw headline numbers (near-perfect ARC-AGI-3, ExploitBench, and SRE Bench scores) gave way within about a day to a more skeptical mood, as commentators amplified reporting on the quietly revised metrics. Reddit threads devoted extensive debate to what commenters called 'benchmaxxing' - whether OpenAI's proprietary-harness ARC-AGI-3 score or ARC Prize's neutral-harness score is the fairer measure of real-world agent capability - with accusations of misleading marketing sitting alongside defenders of harness-augmented scoring as a legitimate capability signal. YouTube's coverage landed calmer: analysts offered measured takes defending the benchmarks' underlying rigor while acknowledging the real hallucination gains were more modest than headline figures implied, and floated architectural theories for the score jumps - a tone distinct from the hype-then-backlash swing on X and Reddit. Even OpenAI's rollout mechanics fed the skepticism - the launch post was delayed almost two hours past its scheduled time, and the company later apologized and offered compensation to paid subscribers who lacked day-one access[10].

Historical Context

2024-12
OpenAI's o3 model faced a prior benchmark-transparency controversy when its claimed 25% FrontierMath score could not be reproduced independently, with Epoch AI's own testing finding closer to 10%.
2025-01
It emerged that OpenAI had secretly funded and had advance access to the FrontierMath benchmark dataset before o3's launch, without disclosing the arrangement to the academic contributors involved.
2026-05
GPT-5.5 was the subject of an industry-wide 'fake thinking' controversy after users found its premium Extended Thinking mode was silently downgraded for many queries.
2026-09-03
GPT-6 Astra launched in limited preview with a delayed, messy blog-post rollout - planned for 2 p.m. ET but not widely viewable until nearly two hours later - followed by the benchmark-revision controversy.

Power Map

Key Players
Subject

GPT-6 Astra launch and benchmark manipulation controversy

OP

OpenAI

Model developer and publisher of the disputed benchmark figures; controls the blog post, system card, and evaluation harness used to generate favorable scores; defended the post-launch edits to press as routine verification.

AR

ARC Prize Foundation (Francois Chollet, Mike Knoop)

Independent creators of the ARC-AGI-3 benchmark; ran Astra under their own neutral harness (62.7%) versus OpenAI's Provider Adapter (99.9%), and publicly declined to endorse OpenAI's AGI framing of the result.

AR

Artificial Analysis

Independent benchmarking firm whose Intelligence Index and hallucination testing contradicted OpenAI's self-reported gains, placing Astra roughly level with its predecessor and behind a rival model.

ST

Stanford researchers (Anka Reuel and Mike Hardy)

Academic critics who reviewed Astra's system card and flagged inadequate documentation of internal benchmarks such as hallucination rate.

AN

Anthropic

Competitor whose Claude Fable 5.1 model figures shifted within OpenAI's own comparison charts and reportedly outperforms Astra on independent indices.

Fact Check

10 cited
  1. [1] OpenAI Astra ARC-AGI-3 harness: 62.7% vs 99.9% benchmark revisions
  2. [2] Benchmarks disagree on GPT-6 Astra but its human-beating efficiency on ARC-AGI-3 pulls Chollet's AGI forecast forward
  3. [3] OpenAI quietly boosts some of Astra's evaluation metrics amid rare delay in publication of the model blog post announcement
  4. [4] OpenAI's Astra, GPT-6, and Greg Brockman on the AGI era
  5. [5] Benchmarking GPT-6 Astra
  6. [6] GPT-6 Astra vs 5 Sol for hallucination, omniscience, accuracy (2026)
  7. [7] OpenAI's o3 AI model scores lower on a benchmark than the company initially implied
  8. [8] We made a mistake in not being more transparent: OpenAI secretly accessed benchmark
  9. [9] GPT-5.5 fake thinking: silent downgrades AI controversy
  10. [10] GPT-6 Astra arrives with major gains, staged access, and new questions about its benchmarks

Source Articles

Top 5

THE SIGNAL.

Analysts

Pushed back directly on OpenAI's suggestion that the ARC-AGI-3 result is evidence of AGI, stating the organization is not making that claim.

Mike Knoop
Co-founder, ARC Prize Foundation

Argued Astra shows genuine internal symbolic-modeling capability that previously required an external harness, and moved up his own AGI timeline forecast in response.

Francois Chollet
Co-creator of the ARC-AGI benchmark series

Criticized Astra's system card for insufficient methodological transparency around key internal evaluations, suggesting the metric edits conveniently benefited OpenAI's marketing narrative.

Anka Reuel and Mike Hardy
Stanford researchers

Offered a more charitable read, framing the metric changes as typical of last-minute launch logistics rather than deliberate manipulation.

Vincent Sunn Chen
Snorkel AI

Suggested that Astra's results may mark the arrival of AGI, a framing the benchmark's own creators declined to endorse.

Greg Brockman
President, OpenAI
The Crowd

OpenAI quietly altered the performance metrics it reported for GPT-6-Astra in ways that favored the new model, and it continues to change others post-launch. Good reporting here from @EmilyForlini

@@jeremyakahn204

Interesting report from FORTUNE: @OpenAI quietly changed several GPT-6 Astra benchmark results around the model's launch, with some changes making Astra look better and competing models worse. One of the biggest examples: Astra's reported hallucination rate went from 4.2% to 2%, before later being changed back to 4.2%. Anthropic Fable 5.1's FrontierMath score also moved from 87.8% to 78%, then back up to 83%. This happened during the strange delay of the Astra launch blog, which was initially published, pulled, and then republished with different numbers. OpenAI says the delay was unrelated to the benchmarks and that eval results can shift depending on the exact checkpoint, harness and reasoning configuration.

@@mark_k274

🚨GPT 6 Astra benchmarks are absolutely ridiculous • ARC AGI 3: 98.6% • ExploitBench: 100% • SRE Bench: 99.2% • AutomationBench: 41.4% What the hell did OpenAI cook up here ??

@@LuminaBench2497

Gpt 6 astra benchmarks

@u/CounterReady47742600
Broadcast
GPT 6 Astra, so good even OpenAI are worried

GPT 6 Astra, so good even OpenAI are worried

GPT-6 Astra Just Went CRITICAL...

GPT-6 Astra Just Went CRITICAL...

GPT-6 Astra Benchmarks: What the AI Industry Won't Tell You

GPT-6 Astra Benchmarks: What the AI Industry Won't Tell You

GPT-6 Astra launch and benchmark manipulation controversy — AI News | Agentic Brew