Welcome to the AGI Era - Or Not
Greg Brockman didn't just announce a model, he announced an era. He called Astra a 'generational leap' and reframed OpenAI's own AGI threshold as 'a mission concept or spiritual concept' [1]. That softening matters: loosening the definition the same week the company claims to have crossed it lets OpenAI claim the milestone rhetorically without inviting the scrutiny a formal declaration would carry.
The benchmark numbers are genuinely striking - Astra saturates FrontierMath Tier 4 at 98%, ARC-AGI-3 at 99.9%, and ExploitBench at 100% [2]- but Gary Marcus pushed back directly on conflating benchmark saturation with AGI: "Success on ARC-AGI is great and impressive, but not - despite the name of the task - proof of AGI" [3]. He also flagged that skeptics, unlike enthusiastic early testers, weren't given advance access to evaluate the model before launch.
The gap between OpenAI's framing and Marcus's skepticism, more than the benchmark scores themselves, is the real story of this launch.


