Google Gemini 3.5 Transcribe launch
TECH

Google Gemini 3.5 Transcribe launch

33+
Signals

Strategic Overview

  • 01.
    Google introduced Gemini 3.5 Transcribe, its most precise speech-to-text model yet, converting raw audio into polished, formatted text across 85+ languages while removing filler words and handling mid-sentence self-corrections.
  • 02.
    The launch ships two distinct model IDs - gemini-3-5-transcribe for pre-recorded audio via the Interactions API with speaker attribution and word-level timestamps, and gemini-3-5-transcribe-live for real-time streaming via the Live API.
  • 03.
    It is already rolling out across Search Live, Gemini Live, Docs, Keep, Gmail, the Gemini app, and Gboard's Rambler dictation feature, and is available to developers via the Gemini API in AI Studio and Antigravity.
  • 04.
    Independent benchmarking from Artificial Analysis ranked it 5th overall on word error rate at 2.6%, with Google citing a 70% improvement in time-to-final-transcription over its predecessor, Chirp 3.

An LLM Wearing a Microphone

Gemini 3.5 Transcribe's most consequential decision is architectural, not cosmetic: it's built on an LLM rather than classical automatic speech recognition, which transcribes audio phoneme-by-phoneme with no real understanding of what's being said. That's why it can handle a correction like "let's meet Tuesday-no, Wednesday" by rewriting the final transcript to just "Wednesday" instead of dutifully capturing both words [1]. The same LLM foundation lets Google position it as understanding intent rather than sound - it removes filler words, auto-formats dates and lists, and adapts to custom vocabulary such as meeting participants' names that a generic model would never recognize on its own. In Google Antigravity specifically, the model pairs screen context and chat history, with permission, to nail file names, agent thoughts, and active documents that would trip up a context-blind transcriber [2].

Fastest and Cheapest, Not Most Accurate - and That's the Pitch

Google isn't claiming Gemini 3.5 Transcribe is the single best transcription model available - it isn't. Independent benchmarking from Artificial Analysis ranks it 5th overall on the AA-WER leaderboard at a 2.6% word error rate for pre-recorded audio [3]- behind pricier specialists like ElevenLabs Scribe v2, which edges it out on raw accuracy alone. What Google is selling instead is the overall package: at roughly $0.005 per minute blended for pre-recorded transcription [4], it undercuts ElevenLabs on price by a wide margin while outrunning rivals like OpenAI's GPT Transcribe on speed - and Google cites a 70% improvement in time-to-final-transcription over its own prior-generation Chirp 3 model [4]. That price-speed-accuracy balance is aimed squarely at existing Google Cloud customers, giving them a reason to consolidate voice workloads onto Gemini rather than routing them to OpenAI's Whisper or dedicated vendors like Deepgram.

The Quiet Keyboard Takeover

The more consequential rollout may not be the API at all. Gemini 3.5 Transcribe already powers Rambler, the dictation feature built into Gboard on Android, turning rambling spoken thoughts into well-formatted text with voice-based editing [5]. It's rolling out simultaneously across Search Live, Gemini Live, Docs, Keep, Gmail, and the Gemini app, and Google has confirmed it's coming to Chrome next, which would extend AI-native dictation into any text field on the open web rather than just Google's own apps [6]. That's a distribution advantage no standalone transcription API can match - Google doesn't need developers to adopt Gemini 3.5 Transcribe for it to reach hundreds of millions of keyboards.

What Early Testers Are Already Pushing Back On

Early reaction across developer and enthusiast communities has skewed positive, with testers highlighting accuracy on notoriously hard cases like Finnish and mid-sentence switching between two languages in the same clip. The reception hasn't been uncritical, though: some testers report that dedicated transcription specialist AssemblyAI still outperforms Gemini 3.5 Transcribe in their own side-by-side comparisons, and there's an open question about feature parity between the two launched variants - the pre-recorded model documents speaker attribution for up to three speakers, but the Live model's own materials don't mention the feature, a gap that's already showing up as a top request in early community feedback. There's also some confusion about Google's own messaging: product marketing pages describe the launch as a public preview, even though Google's own Gemini API changelog already lists the models as generally available (GA), and a few commenters have joked about Google's model-naming conventions, noting the absence of a 'Gemini 3.5 Pro' release.

Historical Context

2025
Chirp 3 was Google's previous-generation transcription model, which Gemini 3.5 Transcribe now replaces with an improved word error rate and a 70% faster time-to-final-transcription.
2026-05-19
Gemini 3.5 Transcribe capabilities were first unveiled at Google I/O before entering public preview in August 2026.
2026-08-12
Google unveiled the Pixel 11 lineup alongside expanded Gemini features, setting up the Gboard Rambler dictation feature that Gemini 3.5 Transcribe now powers.
2026-08-26
Gemini 3.5 Transcribe officially entered public preview via the Gemini API, AI Studio, and Antigravity, and rolled out across Search Live, Gemini Live, Docs, Keep, Gmail, the Gemini app, and Gboard.

Power Map

Key Players
Subject

Google Gemini 3.5 Transcribe launch

GO

Google / Google DeepMind

Developer and publisher of Gemini 3.5 Transcribe; integrates it across first-party consumer products (Search Live, Gemini Live, Docs, Keep, Gmail, Gboard's Rambler) and developer platforms (Gemini API, AI Studio, Antigravity).

IN

IntelliTek Health

Healthcare technology company integrating Gemini 3.5 Transcribe for real-time clinical transcription across primary care and medical specialties, citing accuracy and multi-region compliance benefits.

LI

Lingopal

Real-time translation and broadcast company integrating the model into its speech recognition layer for automatic speaker-language detection and faster multilingual broadcast routing.

OP

OpenAI (Whisper) and dedicated transcription platforms (e.g. Deepgram)

Competitors facing new pressure as Google bundles competitive transcription pricing and performance into the Gemini platform, giving Google Cloud customers a reason to consolidate voice workloads rather than route them elsewhere.

VE

Vercel (AI Gateway)

Third-party developer platform that added Gemini 3.5 Transcribe to its AI Gateway, passing through provider pricing with no markup, extending the model's reach beyond Google's own tooling.

Fact Check

6 cited
  1. [1] Gemini 3.5 Transcribe
  2. [2] Google introduces Gemini 3.5 Transcribe
  3. [3] Gemini 3.5 Transcribe Benchmark and Competitive Analysis
  4. [4] Gemini 3.5 Transcribe: Intelligent Transcription
  5. [5] Gemini 3.5 Transcribe rolls out across Google apps
  6. [6] Gemini Behind Pixel 11's Rambler Is Coming to Chrome

Source Articles

Top 5

THE SIGNAL.

Analysts

Ranked the standard Gemini 3.5 Transcribe model 5th overall on its AA-WER transcription benchmark at a 2.6% word error rate, and measured a 70% improvement in time-to-final-transcription versus Chirp 3.

Artificial Analysis
Independent AI benchmarking organization

By integrating Gemini 3.5 Transcribe, the company says it transforms real-time clinical transcription across primary care and medical specialties, reducing documentation time so providers can focus on patient care.

IntelliTek Health
Healthcare technology integration partner

Says Gemini is consistently among its top model selections and is excited to integrate Gemini 3.5 Transcribe into its speech recognition layer to automatically detect speaker languages and route multilingual broadcast audio faster.

Lingopal
Real-time broadcast translation provider
The Crowd

We're introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet It turns audio into precise transcription in 85+ languages, removing filler words like "ums" and "ahs" while handling self-corrections and capturing your intent so you can get things done using it.

@@Google3955

Say hello to Gemini 3.5 Transcribe! - Build apps that understand user speech / intent, even w/ multiple speakers! - Auto-detection of 85+ languages out of the box - Custom vocab adaptation for specialized jargon... SGTM:) API available now in @GoogleAIStudio and Gemini.

@@sundarpichai2931

Earlier today we introduced Gemini 3.5 Transcribe, our latest text-to-speech model. But, what does this actually mean for your projects? We built this app in @GoogleAIStudio to demonstrate just how much smarter 3.5 Transcribe is. When streaming live audio simultaneously through it.

@@googledevs176

Introducing Gemini 3.5 Transcribe

@u/Stoneonn124
Broadcast
How to build with Gemini 3.5 Transcribe

How to build with Gemini 3.5 Transcribe

Gemini 3.5 Transcribe: 2.6% WER at $0.005/Minute

Gemini 3.5 Transcribe: 2.6% WER at $0.005/Minute

Google's latest speech-to-text Gemini model offers a platter of new features

Google's latest speech-to-text Gemini model offers a platter of new features

Google Gemini 3.5 Transcribe launch — AI News | Agentic Brew