xAI Releases Grok 4.6
TECH

xAI Releases Grok 4.6

87+
Signals

Strategic Overview

  • 01.
    SpaceXAI released Grok 4.6 on August 12, 2026 - a frontier model built for long-running agents, coding, and knowledge work, with a 500,000-token context window and a February 1, 2026 knowledge cutoff.
  • 02.
    The model ships at $2 per million input tokens and $6 per million output tokens for prompts under 200K tokens (rising to $4/$12 above that), undercutting GPT-5.6 Sol's $5/$30 pricing, and adds a new xhigh reasoning-effort tier.
  • 03.
    Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index - a 5-point jump over Grok 4.5 - tying GPT-5.6 Sol but trailing Claude Opus 5 (63) and Claude Fable 5 (62); it also climbed from #13 to #7 on the Code Arena WebDev leaderboard, landing within single digits of GPT-5.6 Sol xHigh and Claude Fable 5.
  • 04.
    The model is live today across the xAI API, Grok Build, Cursor on all plans (with double usage for the first week), Devin Desktop/CLI, and routing platforms OpenRouter, Vercel, and Cloudflare - but there is no open-weights release or self-hosting option.

The Real Story Isn't Performance - It's Margin Compression

Grok 4.6 ships at $2 per million input tokens and $6 per million output tokens below the 200K-token threshold, against $5/$30 for GPT-5.6 Sol and $5/$25 for Claude Opus 5 [1][2]. That is not a rounding difference - investor Gavin Baker quantified it as roughly Claude Fable 5 Max's performance at an 85 percent discount, 80 percent cheaper on input tokens and 88 percent cheaper on output [3]. The pressure lands squarely on OpenAI and Anthropic's inference margins in the run-up to their planned IPOs, and market commentary has already flagged the knock-on risk to Microsoft's equity stakes in both companies as cost-to-outcome becomes the real competitive battleground rather than raw intelligence scores [3].

Grok 4.6 Didn't Get Bigger - It Got Better Trained

It is tempting to read a 5-point jump on the Artificial Analysis Intelligence Index as evidence xAI built a bigger model, but the gain is credited to a longer supplemental training run and expanded reinforcement learning aimed at coding and knowledge work, not new scale [12]. Grok 4.5's underlying V9 foundation was reported at 1.5 trillion parameters [4], and nothing in xAI's own account of Grok 4.6 suggests that changed. Musk himself confirmed in a reply on X that Grok 4.7 is already in training and "significantly better than 4.6," expected in three to four weeks, with a large amount of SpaceX company data now being folded into supplemental training. Grok 4.6 reads less like the leap and more like the bridge.

The Metric That Matters More Than the Leaderboard: Turns Per Task

Grok 4.6's headline Intelligence Index score (61) just ties GPT-5.6 Sol, but the more consequential number sits in the GDPval-AA v2 agentic benchmark: Grok 4.6 resolves tasks in roughly 53 turns and 0.5 billion input tokens, versus about 103 turns and 2.0 billion input tokens for Claude Opus 5 at its max setting [1]. For anyone running agents in production, turn count and token burn compound directly into latency and cost per completed task - a second axis of competitiveness that a single leaderboard score hides. It is the efficiency story, not the raw score, that explains why teams are routing high-volume background agent work to Grok 4.6 while reserving Opus 5 or GPT-5.6 Sol for architectural design and final review [1].

Cheaper Per Token Doesn't Mean Cheaper Per Task - and Not Everyone Is Convinced

Vals Index testing shows Grok 4.6's accuracy climbing to 71.82 percent from Grok 4.5's 65.30 percent, but the practical cost per test actually rose slightly, to $1.61 from $1.25, and latency nearly doubled, from about 388 seconds to 754 seconds [7]. That is the quiet catch behind the price-per-token headline: a model that uses more turns and more tokens per task can erase its own sticker-price advantage. Reception outside the official benchmarks reflects that tension - enthusiasm about the price-performance jump sits alongside real skepticism, particularly among developers comparing hands-on coding results, about whether the benchmark gains translate into everyday task quality or are tuned to the tests themselves. Analysts have also noted that xAI's accelerated release tempo, with Grok 4.7 already queued for weeks out, raises its own trust question: faster iteration widens the window for mistakes to ship before they are caught [6].

Historical Context

2025-07
Grok 4, a multimodal reasoning model, was released.
2025-11
Grok 4.1 launched with improvements to real-world conversation, instruction following, and interaction quality.
2026-04-30
Grok 4.3 shipped as the public flagship model with a 1M-token context window.
2026-07-09
Grok 4.5, xAI's first model built specifically for coding and agentic work, was publicly released, built on a 1.5T-parameter V9 foundation trained with Cursor session data.
2026-07-28
Musk signaled Grok 4.6 would arrive in about two weeks, with Grok 4.7 to follow roughly a month later.
2026-08-12
Grok 4.6 was officially released, gaining 5 points over Grok 4.5 on the Artificial Analysis Intelligence Index just over a month after Grok 4.5's launch.

Power Map

Key Players
Subject

xAI Releases Grok 4.6

XA

xAI / SpaceXAI (Elon Musk)

Developer and publisher of Grok 4.6; Musk publicly framed the release as leapfrogging rivals on cost-efficiency and signaled an aggressive iteration cadence toward Grok 4.7.

OP

OpenAI (GPT-5.6 Sol)

Primary benchmark rival; GPT-5.6 Sol ties Grok 4.6 on the Artificial Analysis Intelligence Index but costs significantly more per token and offers a larger 1M-token context window.

AN

Anthropic (Claude Opus 5 / Claude Fable 5)

Higher-scoring rivals on the Intelligence Index and top-ranked on the GDPval-AA v2 agentic benchmark, but at much higher per-token cost than Grok 4.6.

CU

Cursor

Coding tool that shipped Grok 4.6 immediately on all plans with promotional double usage for the first week.

DE

Devin (Cognition)

Agentic coding platform that integrated Grok 4.6 into Devin Desktop and CLI, highlighting its code-exploration and root-cause-analysis strength.

AR

Artificial Analysis

Independent benchmarking organization that published the headline Intelligence Index comparison placing Grok 4.6 at 61, tied with GPT-5.6 Sol.

Fact Check

12 cited
  1. [1] Grok 4.6: Benchmarks and Analysis
  2. [2] SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price
  3. [3] SpaceX just unveiled Grok 4.6
  4. [4] Grok 4.5 leak: xAI's 1.5T V9 model in private beta
  5. [5] SpaceXAI Releases Grok 4.6
  6. [6] Musk Signals Rapid Grok Rollout: 4.6 in Two Weeks, 4.7 a Month Later
  7. [7] Grok 4.6 - Vals Index
  8. [8] Grok 4.6
  9. [9] Grok 4.6 in Cursor
  10. [10] Grok 4.6 in Devin
  11. [11] Grok 4.6 on OpenRouter
  12. [12] SpaceXAI Releases Grok 4.6

Source Articles

Top 5

THE SIGNAL.

Analysts

Claimed Grok 4.6 is the best model in the world when weighing intelligence, speed, and cost together, and signaled xAI's intent to keep shipping updates rapidly.

Elon Musk
CEO, xAI/SpaceXAI

Framed Grok 4.6 as matching Claude Fable 5 Max's performance at a steep discount, quantifying the input and output token cost gap.

Gavin Baker
Investor

Assessed Grok 4.6 as rejoining the intelligence frontier with standout agentic performance relative to its cost.

Artificial Analysis
Independent AI benchmarking organization

Noted the unusually fast release cadence and speculated Grok could become the fastest-iterating major model if xAI sustains the pace.

Sarah Chen
AI analyst
The Crowd

Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.

@@SpaceXAI29609

@beffjezos Grok 4.7 is significantly better than 4.6 and should be ready in 3 to 4 weeks. Initial training is complete and now we're adding a massive amount of SpaceX company data in supplemental training. This will be something special.

@@elonmusk17322

Grok 4.6 is here! This release by @SpaceXAI just landed in the Code Arena: WebDev at #7 with 1618 pts. Grok 4.6 (High) is a big jump from Grok 4.5, which sits at #13 with 1553 pts. It's now on par with GPT-5.6 Sol xHigh (1622 pts) and Claude Fable 5 (1627 pts). All three currently land in the #5-7 rank range with only 4-9 pts of separation. Expect the picture to sharpen as more votes roll in and CIs tighten. Congrats to @elonmusk and the @SpaceXAI @grok team on this release!

@@arena1544

Grok 4.6 is an equivalent to Sol 5.6 according to artificial analysis arena

@u/Snoo26837794
Broadcast
Design Challenge with Grok 4.6 in Cursor

Design Challenge with Grok 4.6 in Cursor

GROK 4.6 IS HERE ( TESTING LIVE )

GROK 4.6 IS HERE ( TESTING LIVE )

Grok 4.6 Just Changed the AI Race… It Tied GPT-5.6

Grok 4.6 Just Changed the AI Race… It Tied GPT-5.6

xAI Releases Grok 4.6 — AI News | Agentic Brew