The Benchmark You Won't Find in xAI's Launch Post
xAI's own launch materials for Grok 4.6 lean heavily on favorable framing: the model reportedly sits on the Pareto frontier of performance and efficiency on the WANDR benchmark, matching Claude Fable 5's results at over 60% lower cost[1], and Artificial Analysis credits it with a GDPval-AA v2 Elo of 1753, ahead of Claude Fable 5's 1741 and GPT-5.6 Sol's 1728[2]. But the same benchmarking outlet's fuller writeup complicates that lead: on Terminal-Bench v3.0, Grok 4.6 scored just 26%, well behind GPT-5.6 Sol's 34.6% and Claude Fable 5's 34.1%[2], and a separate independent review found a non-hallucination rate of only 65.7% - a figure that never appeared anywhere in xAI's own ten-row evaluation table[3]. The model is also less new than the version bump implies: it keeps the same 1.5 trillion-parameter V9 base as Grok 4.5, with its gains coming from a longer supplemental training run and refined SFT/RL rather than a larger architecture[4]. Reddit users citing Musk's own comments described 4.6 as essentially Grok 4.5 with more reinforcement learning applied, not a larger model.
Real-world testing pushes the picture further from the launch narrative. One practitioner cited in the independent review placed Grok 4.6's coding ability "around Opus 4.8 in real world use, but clearly below Opus 5"[3]. That doesn't erase Grok 4.6's genuine efficiency advantage - Artificial Analysis measured it resolving long-horizon agentic tasks in roughly 53 turns and 0.5 billion input tokens on average, versus about 103 turns and 2.0 billion input tokens for Claude Opus 5 in max mode[2]- but it does mean the "objectively #1" framing xAI used at launch rests on selectively chosen benchmarks rather than a clean sweep across the board.


