xAI 发布 Grok 4.6
TECH

xAI 发布 Grok 4.6

87+
Signals

战略概览

  • 01.
    SpaceXAI 于2026年8月12日发布了 Grok 4.6——这是一款面向长期运行的智能体、编程和知识工作的前沿模型,具备50万 token 的上下文窗口,知识截止日期为2026年2月1日。
  • 02.
    该模型对低于20万 token 的提示按每百万输入 token 2美元、每百万输出 token 6美元定价(超过此阈值则升至4美元/12美元),低于 GPT-5.6 Sol 的5美元/30美元定价,并新增了 xhigh 推理努力层级。
  • 03.
    Grok 4.6 在 Artificial Analysis Intelligence Index 上得分为61,比 Grok 4.5 提高了5分,与 GPT-5.6 Sol 持平,但落后于 Claude Opus 5(63分)和 Claude Fable 5(62分);在 Code Arena WebDev 排行榜上,其排名也从第13位升至第7位,与 GPT-5.6 Sol xHigh 和 Claude Fable 5 的差距已缩小至个位数。
  • 04.
    该模型今日已在 xAI API、Grok Build、所有计划下的 Cursor、Devin Desktop/CLI,以及 OpenRouter、Vercel 和 Cloudflare 等路由平台上线——但未发布开源权重版本,也不支持本地部署。

真正的故事并非性能,而是利润率压缩

Grok 4.6 在20万 token 阈值以下的定价为每百万输入 token 2美元、每百万输出 token 6美元,而 GPT-5.6 Sol 为5美元/30美元,Claude Opus 5 为5美元/25美元[1][2]。这并非微小的价格差异——投资者 Gavin Baker 将其量化为以约85%的折扣提供接近 Claude Fable 5 Max 的性能,输入 token 便宜80%,输出 token 便宜88%[3]。这一压力直接冲击了 OpenAI 和 Anthropic 在计划 IPO 前的推理利润率,市场评论已指出,随着成本效益比成为真正的竞争战场而非原始智能得分,微软在这两家公司中的股权面临连锁风险[3]

Grok 4.6 并没有变大,而是训练得更好了

人们很容易将 Artificial Analysis Intelligence Index 上5分的跃升解读为 xAI 构建了更大的模型,但这一提升归功于更长的补充训练周期和针对编码与知识工作的强化学习扩展,而非规模扩大[12]。据报道,Grok 4.5 的底层 V9 基础模型参数量为1.5万亿[4],而 xAI 对 Grok 4.6 的描述中并未表明这一基础发生了变化。马斯克本人在 X 平台的一条回复中确认,Grok 4.7 已进入训练阶段,且“显著优于4.6”,预计三到四周内发布,大量 SpaceX 公司数据现已纳入补充训练。Grok 4.6 更像是承前启后的桥梁,而非飞跃。

比排行榜更重要的指标:每项任务的交互轮次

Grok 4.6 的 Intelligence Index 得分(61分)仅与 GPT-5.6 Sol 持平,但更具决定性的数字出现在 GDPval-AA v2 智能体基准测试中:Grok 4.6 完成任务平均需约53轮交互和5亿输入 token,而 Claude Opus 5 在最高设置下需约103轮和20亿输入 token[1]。对于生产环境中运行智能体的团队而言,交互轮次和 token 消耗直接转化为延迟和每项完成任务的成本——这是单个排行榜分数所掩盖的竞争第二维度。正是效率优势,而非原始得分,解释了为何团队将高频率后台智能体工作路由至 Grok 4.6,而将 Opus 5 或 GPT-5.6 Sol 保留用于架构设计和最终审查[1]

每 token 更便宜不等于每任务更便宜——并非所有人都信服

Vals Index 测试显示,Grok 4.6 的准确率从 Grok 4.5 的65.30%提升至71.82%,但每次测试的实际成本却略有上升,从1.25美元增至1.61美元,延迟也几乎翻倍,从约388秒升至754秒[7]。这揭示了每 token 价格宣传背后的隐忧:若模型每项任务消耗更多轮次和 token,其标价优势可能被完全抵消。官方基准之外的反馈也反映了这种矛盾——尽管人们对价格性能比的跃升表示热情,但许多开发者在对比实际编码结果时仍持真实怀疑态度,不确定基准测试的提升是否转化为日常任务质量的等效改善,或只是针对测试本身进行了优化。分析师还指出,xAI 加速发布节奏(Grok 4.7 已排期数周后发布)本身带来了信任问题:更快的迭代扩大了错误在被发现前就发布的窗口期[6]

历史背景

发布了 Grok 4,一款多模态推理模型。
Grok 4.1 上线,改进了现实对话、指令遵循和交互质量。
Grok 4.3 作为公众旗舰模型发布,具备100万 token 的上下文窗口。
Grok 4.5 正式发布,这是 xAI 首款专为编程和智能体工作设计的模型,基于1.5万亿参数的 V9 基础模型,使用 Cursor 会话数据进行训练。
马斯克预告 Grok 4.6 将在约两周内发布,Grok 4.7 则将在约一个月后跟进。
Grok 4.6 正式发布,在 Artificial Analysis Intelligence Index 上比 Grok 4.5 提升5分,距离 Grok 4.5 发布仅一个多月。

关键关系图

关键玩家
主题

xAI 发布 Grok 4.6

XA

xAI / SpaceXAI (Elon Musk)

Developer and publisher of Grok 4.6; Musk publicly framed the release as leapfrogging rivals on cost-efficiency and signaled an aggressive iteration cadence toward Grok 4.7.

OP

OpenAI (GPT-5.6 Sol)

Primary benchmark rival; GPT-5.6 Sol ties Grok 4.6 on the Artificial Analysis Intelligence Index but costs significantly more per token and offers a larger 1M-token context window.

AN

Anthropic (Claude Opus 5 / Claude Fable 5)

Higher-scoring rivals on the Intelligence Index and top-ranked on the GDPval-AA v2 agentic benchmark, but at much higher per-token cost than Grok 4.6.

CU

Cursor

Coding tool that shipped Grok 4.6 immediately on all plans with promotional double usage for the first week.

DE

Devin (Cognition)

Agentic coding platform that integrated Grok 4.6 into Devin Desktop and CLI, highlighting its code-exploration and root-cause-analysis strength.

AR

Artificial Analysis

Independent benchmarking organization that published the headline Intelligence Index comparison placing Grok 4.6 at 61, tied with GPT-5.6 Sol.

事实来源

12 条引用
  1. [1] Grok 4.6: Benchmarks and Analysis
  2. [2] SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price
  3. [3] SpaceX just unveiled Grok 4.6
  4. [4] Grok 4.5 leak: xAI's 1.5T V9 model in private beta
  5. [5] SpaceXAI Releases Grok 4.6
  6. [6] Musk Signals Rapid Grok Rollout: 4.6 in Two Weeks, 4.7 a Month Later
  7. [7] Grok 4.6 - Vals Index
  8. [8] Grok 4.6
  9. [9] Grok 4.6 in Cursor
  10. [10] Grok 4.6 in Devin
  11. [11] Grok 4.6 on OpenRouter
  12. [12] SpaceXAI Releases Grok 4.6

来源文章

Top 5

THE SIGNAL.

Analysts

声称 Grok 4.6 在综合考量智能、速度和成本后是全球最佳模型,并表明 xAI 将持续快速发布更新。

Elon Musk
首席执行官,xAI/SpaceXAI

将 Grok 4.6 描述为以大幅折扣实现与 Claude Fable 5 Max 相当的性能,并量化了输入与输出 token 的成本差距。

Gavin Baker
投资者

评估认为 Grok 4.6 凭借相对于成本的突出智能体表现,重新跻身智能前沿行列。

Artificial Analysis
独立 AI 基准测试机构

注意到异常快的发布节奏,并推测若 xAI 能维持此速度,Grok 可能成为迭代最快的主流模型。

Sarah Chen
AI 分析师
The Crowd

Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.

@@SpaceXAI29609

@beffjezos Grok 4.7 is significantly better than 4.6 and should be ready in 3 to 4 weeks. Initial training is complete and now we're adding a massive amount of SpaceX company data in supplemental training. This will be something special.

@@elonmusk17322

Grok 4.6 is here! This release by @SpaceXAI just landed in the Code Arena: WebDev at #7 with 1618 pts. Grok 4.6 (High) is a big jump from Grok 4.5, which sits at #13 with 1553 pts. It's now on par with GPT-5.6 Sol xHigh (1622 pts) and Claude Fable 5 (1627 pts). All three currently land in the #5-7 rank range with only 4-9 pts of separation. Expect the picture to sharpen as more votes roll in and CIs tighten. Congrats to @elonmusk and the @SpaceXAI @grok team on this release!

@@arena1544

Grok 4.6 is an equivalent to Sol 5.6 according to artificial analysis arena

@u/Snoo26837794
Broadcast
Design Challenge with Grok 4.6 in Cursor

Design Challenge with Grok 4.6 in Cursor

GROK 4.6 IS HERE ( TESTING LIVE )

GROK 4.6 IS HERE ( TESTING LIVE )

Grok 4.6 Just Changed the AI Race… It Tied GPT-5.6

Grok 4.6 Just Changed the AI Race… It Tied GPT-5.6