The Harness Gap: How a 62.7% Score Became 99.9% - and an AGI Claim ARC Prize Rejected

OpenAI's headline claim that GPT-6 Astra hit 99.9% on the ARC-AGI-3 benchmark looks far less clean once measured against ARC Prize's own numbers. Under OpenAI's proprietary Provider Adapter - a wrapper that preserves the model's opaque reasoning state across requests and compacts long conversations - Astra scored 99.9% at a cost of $18,817. Under ARC Prize's standard, model-agnostic harness at maximum reasoning, the same model scored 62.7% for $26,098[1]. That is not a rounding difference; it is two different tests wearing the same benchmark's name. OpenAI's own leadership leaned into AGI language around the launch - President Greg Brockman said 'I think it's not unreasonable to feel that we are now in the AGI era'[4]- which makes what came next pointed rather than incidental: ARC Prize itself, the benchmark's creator, explicitly declined to endorse that framing, with co-founder Mike Knoop writing that 'we lack evidence to call this AGI yet'[1]. Co-creator Francois Chollet took a more optimistic read, arguing Astra shows genuine internal symbolic-modeling capability that previously required an external harness to unlock, and moved up his own AGI timeline forecast as a result[2]. Both views can be true at once: the model may be genuinely stronger, but the reported number is a harness artifact, not an apples-to-apples capability score, and it remains well short of the AGI framing OpenAI's own president floated.



