The Benchmark That Depends on How Much You're Willing to Pay

OpenAI's launch materials leaned hard on a headline ARC-AGI-3 score in the high nineties, but the benchmark's own authors pushed back within hours. Francois Chollet, co-founder of the ARC Prize Foundation, confirmed Astra is a genuine step-function jump on interactive reasoning problems - but said it scores 66% on the standard harness, and only approaches 100% when run through a costly custom continuous-conversation harness with manual compaction, at roughly $360 per game[1]. ARC Prize's own writeup was blunt that saturating the benchmark would not amount to 'proof of achieving AGI'[1]. OpenAI also touted strong results elsewhere - FrontierMath Tier 4 v2 at 97.6%, GPQA Diamond at 96%, BenchCAD at 95.9%, and DeepSWE v1.1 at 74.1% - figures that circulated widely on X in benchmark roundups like @wallstengine's rather than being foregrounded on OpenAI's own launch page. Independent aggregator Artificial Analysis went further, measuring Astra's broad Intelligence Index at 61 - exactly matching predecessor GPT-5.6 Sol - despite Astra costing roughly 2.5 times as much to run via the API[2]. The skepticism wasn't confined to specialist benchmark shops: on Reddit, self-identified electrical engineers picked apart a PCB/hardware design demo from OpenAI's launch materials, flagging basic design-rule-check failures, a floating ground plane, and silkscreen errors - a community reaction rather than a formal audit, but one that echoes the same pattern of marketing outrunning delivered capability that Chollet and Artificial Analysis flagged in the benchmarks.


