The Benchmark Numbers Have an Asterisk Problem

OpenAI's launch materials put GPT-6 Astra at 99.9% on ARC-AGI-3, up from 7.8% for GPT-5.6 Sol and 30.2% for Anthropic's Opus 5[1], but that figure comes from OpenAI's own stateful adapter harness; independent stateless runs land Astra at roughly 17-63% depending on reasoning tier[1]. The same pattern repeats on ExploitBench, where OpenAI reports a perfect 100% but a contamination-controlled version built from June-August 2026 vulnerabilities puts Astra at 39.0% versus 5.5% for GPT-5.6 Sol[1], still a real jump, just a smaller one than the marketing implies. FrontierMath Tier 4, a benchmark Epoch AI built independently with more than seventy human mathematicians and no OpenAI involvement, is the one headline number that appears to hold up under scrutiny at around 98%[1]. Muddying things further, Astra scores near the top of one third-party model index but ranks behind Anthropic's Claude Fable 5.1 on Artificial Analysis's intelligence index[2], a reminder that state-of-the-art depends heavily on who is measuring and how.



