Grok 4.6 launch: benchmarks, pricing war, and the Cursor acquisition
TECH

Grok 4.6 launch: benchmarks, pricing war, and the Cursor acquisition

36+
Signals

Strategic Overview

  • 01.
    xAI (branded SpaceXAI) released Grok 4.6 on August 12, 2026 as a post-training upgrade of Grok 4.5, built for long-running agentic, coding, research, and visual work, with pricing held flat at Grok 4.5's rate.
  • 02.
    API pricing is $2 per million input tokens and $6 per million output tokens, but cached-input pricing rose from $0.30 to $0.50 per million tokens and requests over 200,000 tokens are billed at double rates across the entire request.
  • 03.
    Benchmark results are mixed: Grok 4.6 leads on GDPval-AA v2 Elo according to one report (though another source places it behind Claude Opus 5 on the same metric) and matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, but scores well below GPT-5.6 Sol and Claude Fable 5 on Terminal-Bench v3.0.
  • 04.
    SpaceX completed its $60 billion all-stock acquisition of Cursor's parent company, Anysphere, on August 14, 2026, two days after Grok 4.6 launched.

The Benchmark You Won't Find in xAI's Launch Post

xAI's own launch materials for Grok 4.6 lean heavily on favorable framing: the model reportedly sits on the Pareto frontier of performance and efficiency on the WANDR benchmark, matching Claude Fable 5's results at over 60% lower cost[1], and Artificial Analysis credits it with a GDPval-AA v2 Elo of 1753, ahead of Claude Fable 5's 1741 and GPT-5.6 Sol's 1728[2]. But the same benchmarking outlet's fuller writeup complicates that lead: on Terminal-Bench v3.0, Grok 4.6 scored just 26%, well behind GPT-5.6 Sol's 34.6% and Claude Fable 5's 34.1%[2], and a separate independent review found a non-hallucination rate of only 65.7% - a figure that never appeared anywhere in xAI's own ten-row evaluation table[3]. The model is also less new than the version bump implies: it keeps the same 1.5 trillion-parameter V9 base as Grok 4.5, with its gains coming from a longer supplemental training run and refined SFT/RL rather than a larger architecture[4]. Reddit users citing Musk's own comments described 4.6 as essentially Grok 4.5 with more reinforcement learning applied, not a larger model.

Real-world testing pushes the picture further from the launch narrative. One practitioner cited in the independent review placed Grok 4.6's coding ability "around Opus 4.8 in real world use, but clearly below Opus 5"[3]. That doesn't erase Grok 4.6's genuine efficiency advantage - Artificial Analysis measured it resolving long-horizon agentic tasks in roughly 53 turns and 0.5 billion input tokens on average, versus about 103 turns and 2.0 billion input tokens for Claude Opus 5 in max mode[2]- but it does mean the "objectively #1" framing xAI used at launch rests on selectively chosen benchmarks rather than a clean sweep across the board.

Cheaper Than Rivals, But Not the Cheapest in the Room

xAI held Grok 4.6's API pricing flat at $2 per million input tokens and $6 per million output tokens, the same rate as Grok 4.5, well below the per-token rates xAI's rivals charge, per company statements at launch. That gap is real, and it's the number investor Gavin Baker pointed to when he argued Grok 4.6 delivers roughly Claude Fable 5 Max-level performance at an 85% discount[5]. But xAI wasn't the only lab moving that day: DeepSeek released V4 Pro within hours of Grok 4.6, pricing it at $0.435 per million input tokens and $0.87 per million output tokens - roughly seven times cheaper than Grok 4.6 on output tokens alone[6], turning what xAI framed as a pricing win into one leg of a broader same-day price war.

Community cost analysis pushes the pricing story past headline rates and into cost-per-completed-task territory: on the DeepSWE benchmark, Grok 4.6 running in its highest-effort mode scored 67 at roughly $5.50 per task, while GPT-5.6 in max mode matched that score for about $0.61 per task, and DeepSeek V4 Pro scored 63 for roughly $0.06 per task - suggesting that once token-hungry reasoning is factored in, Grok 4.6 isn't necessarily the cheapest way to reach a given quality bar despite its lower advertised rate. xAI's fine print adds a further wrinkle for heavy users: once a request crosses 200,000 tokens, the entire request, not just the overage, is billed at double rates of $4 and $12 per million tokens[7], and cached-input pricing rose from $0.30 to $0.50 per million tokens compared with Grok 4.5[1]. The net effect is a market where "cheaper than OpenAI and Anthropic" holds up, but "cheapest, period" does not.

Why Now: The Cursor Deal Closes the Same Week

SpaceX's $60 billion all-stock acquisition of Cursor's parent company, Anysphere, became effective on August 14, 2026, two days after Grok 4.6 shipped[8]. The timing traces back further than the closing date suggests: SpaceX secured an option in April 2026 to either pay $10 billion for a Cursor partnership or acquire the company outright for $60 billion, a choice it exercised two months later[9]. Reporting on the original deal points to a specific motivation - xAI's coding tools, including Grok itself, had fallen behind Anthropic's Claude Code and OpenAI's Codex in developer market share[9], making an outright purchase of a leading standalone coding-agent product faster than trying to out-build one from scratch. Cursor brought real scale to the deal: its annualized revenue had reached somewhere between roughly $2.6 billion and $4 billion by mid-2026[9], giving xAI both an enterprise customer base and a large pool of real-world coding training data in a single transaction.

The two companies had already been working together before the ink dried: SpaceXAI and Cursor spent months jointly training a model ahead of the acquisition closing, intended for release inside both Cursor and Grok Build. Community discussion on Reddit adds a texture no official announcement includes - a recurring theory among some users that Cursor's engineering team, now folded into xAI, is behind some of Grok 4.6's agentic-coding gains, on the reasoning that much of xAI's original founding team has since departed. Whatever the internal mechanics, the practical result is visible in the product itself: Grok 4.6 launched with double included usage inside Cursor and Grok Build for its first week[11], a promotion that only makes commercial sense once the model and the coding tool are understood as two parts of the same company rather than a vendor relationship.

The Tempo Play: Why Grok 4.7 Is Already Coming

Grok 4.6 is not positioned as a terminal release. Musk had already signaled a rollout cadence weeks before launch - Grok 4.6 arriving on schedule, with Grok 4.7 slated to follow about a month later[10]- and one hands-on reviewer reported Musk describing an already-better internal build shortly after 4.6 shipped. Read one way, that cadence looks like a deliberate strategic choice: shifting the competitive axis away from "who has the single best model on release day" and toward "who controls the fastest iteration loop." Under that lens, Grok 4.6 reads less like a finished flagship and more like a checkpoint - a model released partly to lock in price and distribution across the xAI API, Grok Build, Cursor, OpenRouter, Vercel, and Cloudflare, and newly Perplexity and Perplexity Computer[11], ahead of a capability lead its own maker expects to lose within weeks.

That framing also explains a choice that would otherwise look strange: xAI held pricing completely flat between Grok 4.5 and 4.6 rather than charging a premium for a "new" model[1]. Stable pricing paired with rapid capability churn is the whole point of a tempo strategy - it removes cost as a reason for developers to hesitate on each new checkpoint, while capability increases keep arriving fast enough that competitors are perpetually reacting to the last release rather than the current one. Community speculation ties the next jump to a meaningfully larger model - a rumored 2.1 trillion-parameter Grok 4.7 trained in part on SpaceX's own operational data - which, if accurate, would mark a real departure from the 4.5-to-4.6 pattern of training the same base harder rather than bigger. Either way, every rival lab now has to decide whether to match xAI's release tempo or continue competing on model quality alone, a decision Grok 4.6's flat pricing and same-week Cursor integration were clearly designed to force.

Historical Context

2026-04
SpaceX secured an option to either pay $10 billion for a Cursor partnership or acquire the company outright for $60 billion later in the year.
2026-06-16
SpaceX announced its agreement to acquire Cursor for $60 billion in an all-stock deal.
2026-07
Grok 4.5 launched publicly roughly five weeks before Grok 4.6.
2026-08-12
xAI released Grok 4.6 and DeepSeek released V4 Pro on the same day, triggering a frontier-model price war.
2026-08-14
SpaceX's $60 billion acquisition of Cursor's parent Anysphere became effective.

Power Map

Key Players
Subject

Grok 4.6 launch: benchmarks, pricing war, and the Cursor acquisition

XA

xAI / SpaceXAI (Elon Musk)

Developer and publisher of Grok 4.6; controls pricing and release cadence, and publicly framed the model as leading on intelligence, speed and cost while driving the Cursor acquisition to close a coding-tools gap.

CU

Cursor / Anysphere

AI coding startup acquired by SpaceX for $60B; had been jointly training a model with SpaceXAI ahead of the deal and now brings its enterprise customer base and coding data directly into xAI's stack.

DE

DeepSeek (Liang Wenfeng)

Released DeepSeek V4 Pro the same day as Grok 4.6, undercutting it roughly sevenfold on output pricing and forcing the comparison into a broader frontier price war.

OP

OpenAI / Anthropic

Rival labs whose higher-priced GPT-5.6 Sol and Claude Opus 5 / Claude Fable 5 models are directly undercut on price by Grok 4.6, raising margin-pressure concerns.

PE

Perplexity AI

Added Grok 4.6 to Perplexity and Perplexity Computer shortly after launch, expanding the model's distribution beyond xAI's own products.

Fact Check

11 cited
  1. [1] Introducing Grok 4.6
  2. [2] Grok 4.6: Benchmarks and Analysis
  3. [3] Grok 4.6 Review
  4. [4] SpaceXAI Grok 4.6 Launch, Evals, and Cursor Acquisition
  5. [5] SpaceX Just Unveiled Grok 4.6, Musk Calls It 'Objectively #1' in AI - Anthropic, OpenAI, and This Stock May Be in Trouble
  6. [6] DeepSeek V4 Pro and Grok 4.6 Launch on the Same Day, Igniting a Price War
  7. [7] Grok 4.6 Launch Guide 2026
  8. [8] SpaceX Completes Its $60 Billion Cursor Acquisition
  9. [9] SpaceX Will Buy AI Coding Firm Cursor For $60 Billion
  10. [10] Musk Signals Rapid Grok Rollout: 4.6 in Two Weeks, 4.7 a Month Later
  11. [11] Grok 4.6 in Cursor

Source Articles

Top 5

THE SIGNAL.

Analysts

Claimed Grok 4.6 is "objectively #1 when considering intelligence, speed & cost," positioning it as the top model at the frontier once price is weighed alongside capability.

Elon Musk
CEO, SpaceX / xAI (SpaceXAI)

Argued Grok 4.6 matches Claude Fable 5 Max-level performance at roughly an 85% price discount, framing it as the best price-to-intelligence deal currently at the frontier.

Gavin Baker
Investor

Criticized xAI's launch materials for omitting hallucination-rate data entirely from its evaluation table, reporting a non-hallucination rate of just 65.7% that xAI itself never published.

eesel AI (independent review)
Independent reviewer

Assessed Grok 4.6's real-world coding performance as roughly comparable to an older Claude Opus version, clearly below Claude Opus 5.

Community practitioner (cited in eesel review)
Hands-on tester
The Crowd

Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.

@@SpaceXAI30355

Cursor is now part of @SpaceX. Today, we have officially closed our acquisition. We will join the @SpaceXAI team to help make Grok the world's most useful AI and improve Grok Build, Grok Bot, Grok API, Cursor, and more. SpaceX has built some of the most inspiring and...

@@cursor_ai34854

We've released the model card for Grok 4.6! It goes in-depth on the capabilities of the model across many different evals for coding, engineering, knowledge work, and more. We also discuss pre-deployment safety testing and our safeguard stack.

@@leerob1491

Grok 4.6 is an equivalent to Sol 5.6 according to artificial analysis arena

@u/Snoo26837824
Broadcast
Grok 4.6 IS REALLY GOOD Beating GPT-5.6 Sol, Opus 5, & Kimi k3?! (Fully Tested)

Grok 4.6 IS REALLY GOOD Beating GPT-5.6 Sol, Opus 5, & Kimi k3?! (Fully Tested)

Grok 4.6 Just Shocked OpenAI and Claude

Grok 4.6 Just Shocked OpenAI and Claude

Grok 4.6 Just Changed the AI Race… It Tied GPT-5.6

Grok 4.6 Just Changed the AI Race… It Tied GPT-5.6

Grok 4.6 launch: benchmarks, pricing war, and the Cursor acquisition — AI News | Agentic Brew