The AGI Claim Doesn't Survive Independent Benchmarking
OpenAI markets GPT-6 Astra as its most intelligent and aligned model yet, publishing scores that saturate ARC-AGI-3 at 99.9%, FrontierMath Tier 4 at 98%, and ExploitBench at 100% [1]. OpenAI President Greg Brockman went further, arguing the model's superhuman-speed computer use means 'it's not unreasonable to feel that we are now in the AGI era' [3]. But independent benchmarking firm Artificial Analysis found Astra scores 61 on its Intelligence Index, tying predecessor GPT-5.6 Sol exactly and trailing Anthropic's Claude Fable 5.1 by five points [2]. Forbes contributor Ron Schmelzer raised a related doubt even before the benchmarks landed, questioning whether Astra meets OpenAI's own definition of AGI given how much of its apparent capability leans on surrounding tools rather than the model alone [4]. The gap between OpenAI's headline framing and what a neutral third party actually measured is the clearest sign that Astra's launch was as much a marketing and competitive move as a technical one.


