Gemini 3.8 Live 和 Extended Thinking 发布
TECH

Gemini 3.8 Live 和 Extended Thinking 发布

29+
Signals

战略概览

  • 01.
    谷歌 DeepMind 于 2026 年 9 月 15 日发布了 Gemini 3.8 Live 和 Gemini 3.8 Live Extended Thinking,称这是其迄今为止最先进的实时对话模型。
  • 02.
    Live 版本专为规模化和成本效率设计,具备流畅的对话能力和视觉理解能力,而 Extended Thinking 则针对需要多步推理的高复杂性任务。
  • 03.
    两款模型均可在对话过程中自动检测并切换 97 种支持的语言。
  • 04.
    模型可在后台执行工具和 API 调用,而不会中断实时对话——在确认请求后继续交谈,直到任务完成。
  • 05.
    Extended Thinking 能够边推理边说话,而不是在回应前暂停思考。
  • 06.
    两款模型均属于 Gemini 3 系列,原生支持多模态,可接受音频、图像、视频和文本输入,上下文窗口高达 128K tokens。
  • 07.
    所有由模型生成的音频均带有 SynthID 水印,开发者可通过 Gemini API 和 Google AI Studio 使用这些模型,企业用户可通过 Gemini Enterprise 私有预览版访问,消费者则可通过 Search Live、Gemini Live 以及 Docs、Gmail 和 Keep 等 Workspace 应用使用。

Live 与 Extended Thinking:边思考边说话的技术机制

Gemini 3.8 Live 及其 Extended Thinking 版本基于相同的 Gemini 3 多模态架构构建,支持文本、图像、音频和视频输入,上下文窗口达 128K tokens [2]。两者的区别在于对延迟与深度的处理方式:Live 经过优化以实现规模化和成本效率,结合流畅对话与视觉理解能力;而 Extended Thinking 则设计为在回答时同步进行推理与表达,而非在回应前暂停思考 [1]。两款模型均可在后台启动工具和 API 调用而不中断对话——确认请求后继续聊天,任务在后台完成后将结果整合进来 [1]。两款模型还可在对话中途自动检测并切换 97 种语言 [1]。独立开发者评论指出,其底层架构与 OpenAI 的语音到语音 GPT-Live 模型属于同一家族,可通过类似的 WebSocket API 访问,无需定制集成 [4]

语音 AI 价格战: undercut GPT-Live-1 Astra 与 Grok Voice

谷歌不仅宣称在质量上取胜,更以价格匹配市场。Extended Thinking 每小时音频成本为 3.50 美元,低于 OpenAI 的 GPT-Live-1 Astra(每小时 5.83 美元)和 xAI 的 Grok Voice Think Fast 2.0(每小时 4.80 美元),同时在 Artificial Analysis 的语音到语音质量指数中以 82.6 分领先于对手的 81.5 和 81.3 分 [3]。在代理能力方面同样如此:Extended Thinking 在 tau-Voice 基准测试中以 68.6% 的成绩领先,高于 GPT-Live-1 Astra 的 67.9% 和 Grok Voice 的 56.5%,并在 Sierra 的银行业专用 tau-Voice-banking 排行榜中以 35.1% 领先,远超 GPT-Live-1 Astra 的 32.0% 和 xAI-Realtime 的 16.5% [3]。基础版 Gemini 3.8 Live 定价更低,每小时输入音频仅需 0.84 美元,大幅 undercut 了 GPT-Realtime 2.1 等前代选项 [5]。顶级基准表现与最低每小时费率的结合是一种明确的双线策略:在企业可量化的质量指标上取胜,同时在构建高吞吐量语音代理时赢得单位经济效益。

推理与速度的悖论:为何部分用户更偏爱更快的模型

谷歌自身的营销传递了一个清晰的信息:Live 是快速且便宜的选择,而 Extended Thinking 更智能,隐含假设是更多推理必然更好。基准数据基本支持这一说法——Extended Thinking 在语音到语音质量指数中获得 82.6 分,并在 tau-Voice 和 tau-Voice-banking 测试中领先于 OpenAI 和 xAI 的竞品模型 [3]。但实际用户的发布反馈却呈现出更复杂的图景。社区讨论对语音人格本身存在分歧,一些测试者认为其活泼但令人烦躁,另一些人则觉得它讨人喜欢;更关键的是,有帖子指出 Extended Thinking 在某些交互式排行榜上的表现不如基础版 Live,初步推测是实时对话更看重速度而非深思熟虑,用户并不希望语音代理在回应前明显停顿思考。这种张力——一个被设计为边说边想的模型,却遭遇偏好几乎不思考的用户的现实——或许是此次发布留下的最有趣未解问题:更聪明的语音代理是否真的等同于更好的语音代理?

从 I/O 演示到生产落地:Project Astra 如何演进为 Gemini 3.8 Live

Gemini 3.8 Live 并非凭空出现。谷歌首次在 2024 年 I/O 大会上推出与 Gemini 对话的模式,并配合 viral 的 Project Astra 演示,展示了接近实时的多模态 AI 能力 [6]。一年后,Astra 的低延迟多模态技术被整合进 Google Search、Gemini 应用和开发者工具中,为本次发布的功能奠定了技术基础 [7]。与此前相比,此次发布的变化在于范围:新版本将后台工具执行、97 种语言切换,以及在廉价快速模型与高推理模型之间的真正选择,统一整合进一个生产就绪的 API 和 AI Studio 界面,并集成至 Docs、Gmail 和 Keep 等 Workspace 应用 [1]。然而,功能扩展的代价是,谷歌自身的模型卡承认这些模型仍可能存在基础大模型的通用局限,例如幻觉问题,而在始终在线、执行工具的语音代理中,这一风险比在普通聊天窗口中更为严重 [2]。对于已使用 Gemini 文本模型进行代理工作流的企业而言,拥有一个同家族、定价显著低于竞品语音模型的语音选项,是更具吸引力的商业卖点 [5]

历史背景

谷歌在 2024 年 I/O 大会上宣布推出 Gemini Live,允许用户通过移动应用与 Gemini 对话,同时发布了展示接近实时多模态 AI 的 viral Project Astra 演示。
Project Astra 的低延迟、多模态能力被扩展至 Google Search、Gemini 和开发者工具中,为如今 Gemini 3.8 Live 中的实时语音功能奠定了基础。
Google DeepMind 推出了 Gemini 3.8 Live 和 Gemini 3.8 Live Extended Thinking,这是其最新的实时语音模型,支持后台工具调用和 97 种语言切换。

关键关系图

关键玩家
主题

Gemini 3.8 Live 和 Extended Thinking 发布

GO

Google DeepMind

Developer of Gemini 3.8 Live and Extended Thinking; announced the launch as its best conversational AI, extending the Gemini 3 model family into real-time voice.

OP

OpenAI (GPT-Live-1 Astra)

Direct competitor whose speech-to-speech model was outranked by Gemini 3.8 Live Extended Thinking on the Artificial Analysis Speech-to-Speech Quality Index (82.6 vs 81.5) and priced higher at $5.83/hour vs Google's $3.50/hour.

XA

xAI (Grok Voice / Think Fast 2.0)

Competitor whose Grok Voice Think Fast 2.0 (High) scored 81.3 on the same quality index and costs $4.80/hour, more expensive and lower-scoring than Gemini 3.8 Live Extended Thinking.

AR

Artificial Analysis

Independent benchmarking organization whose Speech to Speech Quality Index and tau-Voice agentic benchmark rankings underpin Google's third-party performance claims for the new models.

事实来源

7 条引用
  1. [1] Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking
  2. [2] Gemini 3.8 Audio Model Card
  3. [3] Google Releases Gemini 3.8 Live-Extended Conversational Model, Claims Better Performance Than Rivals At Lower Price
  4. [4] Gemini 3.8 Live and 3.8 Live Extended Thinking
  5. [5] Google Gemini 3.8 Live Extended Thinking Announcement
  6. [6] Gemini Live announced at Google I/O 2024
  7. [7] Project Astra comes to Google Search, Gemini and developers

来源文章

Top 5

THE SIGNAL.

Analysts

将 Gemini 3.8 Live 视为结构上与 OpenAI 的 GPT-Live 系列相当的模型,指出其可通过类似现有语音到语音产品的 WebSocket 实现进行访问。

Simon Willison
独立 AI 开发者与评论员
The Crowd

Introducing our most advanced Gemini Audio models yet 🗣 Gemini 3.8 Live and 3.8 Live Extended Thinking let you speak, collaborate, and execute tasks seamlessly, meaning conversing with AI just got a lot more natural. So, what's the difference between these two models? Let's

@@GoogleAI2853

Google has released Gemini 3.8 Live, its new Speech to Speech model, with the Extended Thinking (High) variant debuting at #1 on the Artificial Analysis Speech to Speech Index at 82.6, and #1 on our Tau Voice benchmark implementation at 68.6% Gemini 3.8 Live is @GoogleDeepMind's

@@ArtificialAnlys1200

My mind is f*** blown. I built this live insurance claim agent that can see, talk, think and draw in real-time using the new Gemini 3.8 LIVE. Even switched my language to Hindi mid-call and it still worked. VOICE AI can't be more real. Made it 100% open-source.

@@Saboo_Shubham_1009

Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking

@u/gibbonwalker268
Broadcast
What's new in the Gemini Live API

What's new in the Gemini Live API

Gemini 3.8 Live VE y HABLA en tiempo real (y da miedo)

Gemini 3.8 Live VE y HABLA en tiempo real (y da miedo)

【環境最強】リアルタイムAI『Gemini 3.8 Live』が最強コスパで登場!精度も体験も良いので解説

【環境最強】リアルタイムAI『Gemini 3.8 Live』が最強コスパで登場!精度も体験も良いので解説