A Comeback Built on Selective Benchmarks
Google is positioning Gemini 4 Argon as a return to frontier-model leadership after ceding recent benchmark headlines to OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5 [1]. The model leads or ties on 13 of 18 disclosed benchmarks, with its widest margins in enterprise knowledge work: Argon scores 19.6% on Harvey's Legal Agent Benchmark against 5.4% for GPT-6 Astra and 3.8% for Claude Opus 5.5, and 51.3% on AutomationBench against 42.5% and 41.4% for the same two rivals [1]. On real-world software engineering it posts 77.9% on DeepSWE v1.1, again ahead of both competitors [1]. The gaps run the other way on two coding-specific tests - GPT-6 Astra leads FrontierSWE v2 by more than 10 points (65.5% to 55.0%), and Claude Opus 5.5 leads Terminal-Bench 4.0 (66.4% to 57.4%) [1]. That split has not gone unnoticed: Reddit's r/singularity threads accused Google of "benchmaxxing," arguing Argon's strongest scores cluster on the easier end of the suite, and cited a Bloomberg report claiming that employees who have put the model to work found it struggles on certain real-world coding tasks despite the strong headline numbers. Google's messaging leans into breadth rather than any single score - Tulsee Doshi, who leads Gemini products at DeepMind, calls Argon "a well-rounded model that has frontier capabilities across several domains" rather than a benchmark-tuned specialist [2].



