The AGI headline number didn't survive stricter testing

Astra's most viral scores came from stateful, tool-assisted evaluation runs, including a near-perfect ARC-AGI-3 result [6]. But on ExploitBench, a 'perfect' 100% score collapsed to 39% once the benchmark was rebuilt to exclude historical, previously-disclosed vulnerabilities [6]- a strong signal that some of that ceiling reflects memorized exploit data rather than genuinely novel discovery. OpenAI's own framing leaned hard into the opposite conclusion: President Greg Brockman said it was 'not unreasonable' to feel the industry has entered the AGI era [8]. One widely-watched developer review pushed back on that framing with a different measure entirely - an independent, non-benchmark-gamed intelligence index that placed Astra level with its own predecessor and behind a rival lab's model, despite Astra's benchmark sheet looking like a clean sweep.


![Why GPT-6 Astra Drops from 99.9% to 62.7% [Complete Breakdown]](https://img.youtube.com/vi/g5rJ-gkrBcA/mqdefault.jpg)