GPT-6 Astra launch, capabilities, and post-launch quality issues
TECH

GPT-6 Astra launch, capabilities, and post-launch quality issues

33+
Signals

Strategic Overview

  • 01.
    OpenAI launched GPT-6 Astra on September 3, 2026, first to a limited set of organizations before expanding within days to all ChatGPT Plus, Pro, Business, and Enterprise users, the API, and AWS.
  • 02.
    The model is also generally available to all customers in Microsoft Foundry on Azure, with tiered pricing for short- and long-context requests.
  • 03.
    Astra posted state-of-the-art results on multiple reasoning, coding, and computer-use benchmarks, and became the first OpenAI model to cross the company's 'Critical' cybersecurity capability threshold.
  • 04.
    It ships with a 1,050,000-token context window, a 128,000-token max output, and an April 30, 2026 knowledge cutoff.
  • 05.
    Within days of launch, users reported a noticeable drop in output quality; OpenAI diagnosed three separate defects and issued a full usage reset to fix them.
  • 06.
    OpenAI President Greg Brockman suggested the model's capabilities mark the arrival of the 'AGI era.'

The AGI headline number didn't survive stricter testing

The AGI headline number didn't survive stricter testing
ExploitBench score before and after removing memorized historical vulnerabilities.

Astra's most viral scores came from stateful, tool-assisted evaluation runs, including a near-perfect ARC-AGI-3 result [6]. But on ExploitBench, a 'perfect' 100% score collapsed to 39% once the benchmark was rebuilt to exclude historical, previously-disclosed vulnerabilities [6]- a strong signal that some of that ceiling reflects memorized exploit data rather than genuinely novel discovery. OpenAI's own framing leaned hard into the opposite conclusion: President Greg Brockman said it was 'not unreasonable' to feel the industry has entered the AGI era [8]. One widely-watched developer review pushed back on that framing with a different measure entirely - an independent, non-benchmark-gamed intelligence index that placed Astra level with its own predecessor and behind a rival lab's model, despite Astra's benchmark sheet looking like a clean sweep.

First to cross OpenAI's 'Critical' cyber threshold - with less oversight

Astra is the first OpenAI model to cross the company's Preparedness Framework 'Critical' cybersecurity threshold: with appropriate tools and access, it can find previously unknown security flaws and develop novel exploits across well-protected systems largely without step-by-step human guidance [1]. In OpenAI's own authorized-target safety test, Astra stayed within bounds 100% of the time without production safeguards, versus GPT-5.6 Sol overstepping 48% of the time [1]. But Greyhound Research's Sanchit Vir Gogia pushed back on the framing itself: 'Astra's capability did not change between 10 August ... and September 1 ... The testing changed. The model did not' [1]. He also flagged that OpenAI's own safety reporting shows decreased chain-of-thought monitorability compared with the predecessor model, meaning better behavior on paper has come with reduced ability to inspect why the model does what it does [1].

Spatial reasoning made a real jump - just not to human level

Independent of OpenAI's own claims, Cornell and Google DeepMind researcher Yoav Artzi ran early benchmark tests and concluded Astra 'does seem like a step change in spatial reasoning' [7]. On the StationeryBench desk-manipulation benchmark, Astra fully completed 7 of 100 tasks versus zero for Ai2's MolmoAct2, and posted a median progress score of 46/100 against MolmoAct2's 12/100 [7]. That's a genuine gap over the prior generation of models. Even so, Artzi's own REMAP testing found Astra still falls short of human performance on some of the same 3D reasoning scenarios [7]- a useful check on how far 'step change' actually goes.

Three bugs, one reset: the launch's real controversy

Days after launch, users across ChatGPT and Codex reported that Astra's agentic output had visibly degraded - stopping tasks early, claiming work was done when it wasn't, or responding to stale context [4]. OpenAI's Codex product lead Tibo Sottiaux publicly attributed this to three distinct defects: legacy skills written for older models over-triggering and sometimes blocking Astra from checking its own work; a broken opt-in context-management experiment that made the model quit early or reply to stale messages, affecting an estimated 4,000 to 5,000 users before it was disabled; and inference engines serving tail traffic that were misconfigured and measurably degraded output quality [4]. The fix shipped as a full usage reset pushed out before midnight local time [4]. OpenAI's own engineering team walked through the same three causes in a public reset announcement, directly countering community speculation - visible in both Reddit threads and OpenAI staff replies - that the model had quietly been swapped for a cheaper, more compressed version to handle demand.

Rollout: rocky start, quick enterprise expansion

Astra rolled out first to a limited set of organizations on September 3, then expanded within days to all ChatGPT Plus, Pro, Business, and Enterprise plans, plus the API and AWS [1]. OpenAI's own announcement confirms it is live for Pro, Enterprise, and Business Premium users in ChatGPT Work and Codex [2]. Microsoft moved just as fast, making Astra generally available to all Foundry customers on Azure, priced at $10 / $1 / $50 per million input, cached-input, and output tokens for short-context requests in the Standard Global tier, rising to $20 / $2 / $75 for long context [3]. The model ships with a 1.05 million-token context window, a 128,000-token max output, and an April 30, 2026 knowledge cutoff [2].

Historical Context

2026-09-03
GPT-6 Astra initially released to approved/limited organizations, with a rocky rollout that briefly locked out paying ChatGPT subscribers.
2026-09-04
General availability of GPT-6 Astra extended to ChatGPT Plus/Pro/Business/Enterprise users and via API/Azure/AWS.
2026-09-05
OpenAI published developer guidance ('Rethinking skills and prompts for GPT-6 Astra') recommending leaner prompts and fewer guardrails for the new model.
2026-09-06
Users began reporting sustained degraded agentic behavior from Astra (premature turn termination, false completion reports).

Power Map

Key Players
Subject

GPT-6 Astra launch, capabilities, and post-launch quality issues

OP

OpenAI

Developer and publisher of GPT-6 Astra; controls rollout, pricing, safety classification, and post-launch bug fixes/reset.

TI

Tibo Sottiaux

OpenAI Codex product lead; publicly diagnosed the three post-launch quality bugs and announced the fixes and usage reset.

ER

Eric Provencher

OpenAI staffer who authored developer guidance on simplifying skills/prompts for Astra, influencing how enterprise developers adapt their agent tooling.

MI

Microsoft (Azure/Foundry)

Enterprise distribution partner; made Astra generally available in Microsoft Foundry with its own regional pricing tiers, expanding enterprise reach.

YO

Yoav Artzi (Cornell / Google DeepMind)

Independent AI researcher who benchmarked Astra's spatial reasoning and robotics performance (StationeryBench, REMAP), providing outside validation of capability claims.

SA

Sanchit Vir Gogia (Greyhound Research)

Chief analyst providing critical governance commentary, arguing the 'Critical' cyber classification reflects a testing/disclosure change rather than a capability change, and flagging reduced chain-of-thought monitorability.

Fact Check

8 cited
  1. [1] OpenAI launches GPT-6 Astra, its first model to cross a critical cybersecurity threshold
  2. [2] Introducing GPT-6 Astra: the most intelligent and aligned model in the world
  3. [3] GPT-6 Astra: frontier intelligence for work now generally available in Microsoft Foundry
  4. [4] OpenAI defects: GPT-6 Astra's sudden slump
  5. [5] GPT-6 Astra needs leaner prompts and fewer guardrails, OpenAI recommends
  6. [6] GPT-6 Astra Benchmarks Explained
  7. [7] GPT-6 Astra appears to show a 'step change' in spatial reasoning based on early benchmarks
  8. [8] OpenAI launches GPT-6 Astra, and president Greg Brockman says we may be entering the AGI era

Source Articles

Top 5

THE SIGNAL.

Analysts

Praised Astra's spatial reasoning as a major leap forward via independent robotics benchmarks, while noting it still lags human performance in some scenarios. Quote: "Astra does seem like a step change in spatial reasoning. Impressive..."

Yoav Artzi
AI researcher, Cornell University and Google DeepMind

Argues that Astra's 'Critical' cybersecurity classification reflects a change in OpenAI's testing methodology rather than a genuine jump in the model's underlying capability: "Astra's capability did not change between 10 August, when OpenAI said Critical capability could not be ruled out, and September 1, when it said the threshold was met. The testing changed. The model did not."

Sanchit Vir Gogia
Chief analyst, Greyhound Research

Warns that Astra's improved behavior comes with reduced oversight capability compared to its predecessor: "Astra behaves better and watches worse: OpenAI reports decreased chain-of-thought monitorability against Sol."

Sanchit Vir Gogia
Chief analyst, Greyhound Research

Advises developers to strip down legacy scaffolding (skills, AGENTS.md, prompts) built for weaker models since Astra needs less hand-holding: "Astra can figure out what it needs on its own."

Eric Provencher
OpenAI developer relations

Suggests that Astra could be looked back on as the model marking the beginning of the AGI era: "I think it's not unreasonable to feel that we are now in the AGI era."

Greg Brockman
President, OpenAI

Frames Astra's significance as illustrating a shared industry dilemma of balancing frontier capability gains against the efficacy of safe monitoring mechanisms.

Nick Patience
VP and Practice Lead, Futurum Group
The Crowd

Hi Astra users. A reset and a quick update on quality issues that have been posted around. Working with some of you, we have found and fixed the following issues: - Some skills written for previous models were triggering too often or preventing the model from checking its work.

@@thsottiaux27361

@ChaseLochmiller @OpenAI GPT-6 Astra, trained on ~100K+ NVIDIA Grace Blackwell NVLink72. From ChatGPT to o1 to Astra in 4 years. AGI has arrived. Congratulations @OpenAI team. 400K GPUs coming online next.

@@JensenHuang41784

Reset rolling out to all Codex & ChatGPT Work users! Grateful to everyone who helped us investigate the Astra quality issues and shared examples. We've identified and fixed issues with: > Skills over-triggering or preventing self-checks > Context management causing early stops

@@reach_vb1025

GPT 6 Astra already downgraded?

@u/SteveEricJordan303
Broadcast
Did OpenAI actually build AGI? GPT-6 Astra first look

Did OpenAI actually build AGI? GPT-6 Astra first look

Introducing GPT-6 Astra: the most intelligent and aligned model in the world.

Introducing GPT-6 Astra: the most intelligent and aligned model in the world.

Why GPT-6 Astra Drops from 99.9% to 62.7% [Complete Breakdown]

Why GPT-6 Astra Drops from 99.9% to 62.7% [Complete Breakdown]