Google Gemini 3.5 Transcribe 发布
TECH

Google Gemini 3.5 Transcribe 发布

33+
Signals

战略概览

  • 01.
    Google 推出了 Gemini 3.5 Transcribe,这是其迄今为止最精确的语音转文本模型,可将原始音频转换为经过润色、格式化的文本,支持 85 种以上语言,并能去除填充词,处理句中自我纠正的情况。
  • 02.
    此次发布推出了两个独立的模型 ID —— gemini-3-5-transcribe 通过 Interactions API 支持预录制音频,具备说话人归属识别和词级时间戳;gemini-3-5-transcribe-live 则通过 Live API 支持实时流式传输。
  • 03.
    该功能已在 Search Live、Gemini Live、Docs、Keep、Gmail、Gemini 应用以及 Gboard 的 Rambler 语音输入功能中逐步上线,并可通过 AI Studio 和 Antigravity 中的 Gemini API 向开发者开放。
  • 04.
    根据 Artificial Analysis 的独立基准测试,其词错误率(WER)总体排名第 5,为 2.6%,Google 表示其最终转录耗时相比前代 Chirp 3 提升了 70%。

戴上麦克风的大语言模型

Gemini 3.5 Transcribe 最具影响力的决策是架构层面而非外观上的:它基于大语言模型(LLM)构建,而非传统的自动语音识别系统。传统 ASR 系统逐音素转录语音,对所说内容毫无理解。正因如此,当用户说“我们周二见面——不,周三”时,该模型能在最终文稿中直接修正为“周三”,而不是机械地记录两个词 [1]。相同的 LLM 基础使 Google 能将其定位为理解意图而非声音的工具——它能删除填充词,自动格式化日期和列表,并适应自定义词汇,例如会议参与者的姓名,这些是通用模型无法自行识别的内容。在 Google Antigravity 中,该模型在获得许可的前提下结合屏幕上下文和聊天历史,精准识别文件名、代理思路和当前文档,避免上下文盲区导致的识别错误 [2]

最快且最便宜,而非最准确——而这正是卖点

Google 并未声称 Gemini 3.5 Transcribe 是目前最优秀的转录模型——事实也并非如此。Artificial Analysis 的独立基准测试显示,在 AA-WER 榜单上,其预录制音频的词错误率为 2.6%,总体排名第五 [3],落后于 ElevenLabs Scribe v2 等更昂贵的专业模型,后者仅在原始准确性上略胜一筹。Google 所推销的是整体方案:预录制转录综合价格约为每分钟 0.005 美元 [4],远低于 ElevenLabs 的定价,同时在速度上超越 OpenAI 的 GPT Transcribe 等竞争对手——Google 称其相比自家前代 Chirp 3 模型,最终转录耗时缩短了 70% [4]。这种价格、速度与准确性的平衡,明确针对现有 Google Cloud 客户,促使他们将语音工作负载整合至 Gemini,而非继续使用 OpenAI 的 Whisper 或 Deepgram 等专用供应商。

悄然进行的键盘接管

更具深远影响的可能不是 API 本身。Gemini 3.5 Transcribe 已经驱动 Android 上 Gboard 内置的语音输入功能 Rambler,将杂乱的口语转化为格式清晰的文本,并支持语音编辑 [5]。它正在同步登陆 Search Live、Gemini Live、Docs、Keep、Gmail 和 Gemini 应用,Google 还确认其即将登陆 Chrome 浏览器,这意味着 AI 原生语音输入将扩展到整个开放网络的任意文本字段,而不仅限于 Google 自家应用 [6]。这种分发优势是任何独立转录 API 都无法比拟的——Google 无需开发者主动采用 Gemini 3.5 Transcribe,就能让其触达数亿台设备的键盘。

早期测试者已经开始提出哪些质疑

开发者和爱好者社区的初步反应总体积极,测试者特别强调其在芬兰语等公认难题以及同一音频片段中句内双语切换场景下的高准确率。但也有批评声音:一些测试者报告称,在他们的并排对比中,专业转录服务商 AssemblyAI 仍优于 Gemini 3.5 Transcribe;此外,两个已发布版本之间是否存在功能对等性尚存疑问——预录制模型文档标明支持最多三位说话人的身份识别,但 Live 模型资料中并未提及此功能,这一差距已在早期社区反馈中成为最高频的需求之一。关于 Google 自身的信息传达也存在困惑:产品营销页面将其描述为公开预览版,而 Google 自己的 Gemini API 更新日志却已将这些模型列为正式发布(GA)。少数评论者还调侃了 Google 的命名惯例,指出尚未发布“Gemini 3.5 Pro”版本。

历史背景

Chirp 3 是 Google 上一代转录模型,Gemini 3.5 Transcribe 现已取代它,在词错误率方面有所改进,并将最终转录耗时缩短了 70%。
Gemini 3.5 Transcribe 功能首次在 Google I/O 上揭晓,随后于 2026 年 8 月进入公开预览。
Google 在发布 Pixel 11 系列的同时扩展了 Gemini 功能,为 Gboard 的 Rambler 语音输入功能奠定基础,而该功能如今由 Gemini 3.5 Transcribe 驱动。
Gemini 3.5 Transcribe 正式通过 Gemini API、AI Studio 和 Antigravity 进入公开预览,并在 Search Live、Gemini Live、Docs、Keep、Gmail、Gemini 应用和 Gboard 中全面推出。

关键关系图

关键玩家
主题

Google Gemini 3.5 Transcribe 发布

GO

Google / Google DeepMind

Developer and publisher of Gemini 3.5 Transcribe; integrates it across first-party consumer products (Search Live, Gemini Live, Docs, Keep, Gmail, Gboard's Rambler) and developer platforms (Gemini API, AI Studio, Antigravity).

IN

IntelliTek Health

Healthcare technology company integrating Gemini 3.5 Transcribe for real-time clinical transcription across primary care and medical specialties, citing accuracy and multi-region compliance benefits.

LI

Lingopal

Real-time translation and broadcast company integrating the model into its speech recognition layer for automatic speaker-language detection and faster multilingual broadcast routing.

OP

OpenAI (Whisper) and dedicated transcription platforms (e.g. Deepgram)

Competitors facing new pressure as Google bundles competitive transcription pricing and performance into the Gemini platform, giving Google Cloud customers a reason to consolidate voice workloads rather than route them elsewhere.

VE

Vercel (AI Gateway)

Third-party developer platform that added Gemini 3.5 Transcribe to its AI Gateway, passing through provider pricing with no markup, extending the model's reach beyond Google's own tooling.

事实来源

6 条引用
  1. [1] Gemini 3.5 Transcribe
  2. [2] Google introduces Gemini 3.5 Transcribe
  3. [3] Gemini 3.5 Transcribe Benchmark and Competitive Analysis
  4. [4] Gemini 3.5 Transcribe: Intelligent Transcription
  5. [5] Gemini 3.5 Transcribe rolls out across Google apps
  6. [6] Gemini Behind Pixel 11's Rambler Is Coming to Chrome

来源文章

Top 5

THE SIGNAL.

Analysts

在 AA-WER 转录基准测试中,将标准版 Gemini 3.5 Transcribe 模型总体排名第五,词错误率为 2.6%,并测得其相比 Chirp 3 的最终转录耗时提升了 70%。

Artificial Analysis
独立人工智能基准测试机构

通过集成 Gemini 3.5 Transcribe,该公司表示可在初级护理和医学专科中实现临床实时转录的转型,减少文书工作时间,使医护人员能更专注于患者照护。

IntelliTek Health
医疗科技集成合作伙伴

表示 Gemini 一直是其首选模型之一,并期待将 Gemini 3.5 Transcribe 集成至其语音识别层,以更快地自动检测发言人语言并路由多语言广播音频。

Lingopal
实时广播翻译服务提供商
The Crowd

We're introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet It turns audio into precise transcription in 85+ languages, removing filler words like "ums" and "ahs" while handling self-corrections and capturing your intent so you can get things done using it.

@@Google3955

Say hello to Gemini 3.5 Transcribe! - Build apps that understand user speech / intent, even w/ multiple speakers! - Auto-detection of 85+ languages out of the box - Custom vocab adaptation for specialized jargon... SGTM:) API available now in @GoogleAIStudio and Gemini.

@@sundarpichai2931

Earlier today we introduced Gemini 3.5 Transcribe, our latest text-to-speech model. But, what does this actually mean for your projects? We built this app in @GoogleAIStudio to demonstrate just how much smarter 3.5 Transcribe is. When streaming live audio simultaneously through it.

@@googledevs176

Introducing Gemini 3.5 Transcribe

@u/Stoneonn124
Broadcast
How to build with Gemini 3.5 Transcribe

How to build with Gemini 3.5 Transcribe

Gemini 3.5 Transcribe: 2.6% WER at $0.005/Minute

Gemini 3.5 Transcribe: 2.6% WER at $0.005/Minute

Google's latest speech-to-text Gemini model offers a platter of new features

Google's latest speech-to-text Gemini model offers a platter of new features