The Benchmark Gap Between Google's Chart and the Independent Re-Run
Google says Argon leads or ties on 12 to 13 of 18 to 19 disclosed benchmarks against GPT-6 Astra and Claude Opus 5.5, including a self-reported 77.9% on DeepSWE v1.1 [1]. That DeepSWE number, though, is Google's own computation - it doesn't appear on the independent DataCurve leaderboard at all, where the top scorers tie at 74%, one of them Google's own smaller Gemini 3.8 Flash. Other independent evaluations split the difference rather than confirm the lead: Vals AI ranks Argon first of 41 models overall, while Artificial Analysis' Intelligence Index puts it level with GPT-6 Astra at 53, behind Claude Opus 5.5's 60 [2], and ranks it only eighth of 223 models in its price class, with Claude Sonnet 5.5 beating it head-to-head on coding-specific tests. None of this gets settled by argument; it gets settled by running the model on an unfamiliar codebase, which almost nobody outside the Fairwind Program can currently do.


