Google launches Gemini 3.5 Transcribe speech-to-text model
TECH

Google launches Gemini 3.5 Transcribe speech-to-text model

33+
Signals

Strategic Overview

  • 01.
    Google introduced Gemini 3.5 Transcribe, described as its most precise speech-to-text model yet, converting raw audio directly into accurate, polished, formatted text that handles background noise, jargon, and disfluency cleanup.
  • 02.
    The model automatically detects and transcribes over 85 languages, including regional accents and dialects, with no manual language selection required.
  • 03.
    It ships as two separate API surfaces: the Live API (gemini-3.5-transcribe-live) for real-time streaming and the Interactions API (gemini-3.5-transcribe) for pre-recorded, batch audio processing.
  • 04.
    The model is rolling out across first-party Google surfaces including Search Live, Gemini Live, Docs, Keep, Gmail, the Gemini app, and Gboard, already powering the Rambler feature on Pixel 11-series phones and the Gemini app on macOS, with Chrome integration coming soon.

Positioning: A Gemini Model, Not a Cloud Speech Spinoff

The most consequential detail here isn't a feature - it's a product-structure decision. Commentary notes this is the first Google speech-to-text model genuinely sold as a Gemini model, sharing the same API surface, including native function calling, as the rest of the Gemini family, rather than living under the separate Cloud Speech product line [1]. Google's earlier ASR generation, Chirp 3, was announced for Vertex AI as its own distinct product; Gemini 3.5 Transcribe is explicitly framed as a direct successor to Chirp 3, but the successor relationship is architectural as much as generational [2]. Practically, that means a developer already calling Gemini for text or multimodal reasoning can add speech transcription and function-calling-driven voice actions through the same API key and surface, instead of standing up a separate Cloud Speech integration. It reframes transcription as a capability of the flagship model line rather than a bolt-on utility product - a strategy move that mirrors how Google has folded other modalities into Gemini over time.

Two APIs, One Model Family: Live Streaming vs Batch Transcription

Gemini 3.5 Transcribe actually ships as two distinct endpoints tuned for different jobs: the Live API, using the model gemini-3.5-transcribe-live, handles real-time streaming audio, while the Interactions API, using gemini-3.5-transcribe, handles pre-recorded or batch audio [2]. The split shows up directly in the benchmark numbers Google published: streaming transcription averages 4.0% word error rate, while non-streaming (batch) transcription is markedly cleaner at 2.6% average WER, reflecting the tradeoff between speed and the model's ability to revise its output before finalizing [3]. On the latency side, the streaming variant returns a final transcript just 0.40 seconds after speech ends - a figure one analysis frames as the single number that determines whether a voice-agent vendor wins enterprise contracts [4]. That combination - a genuinely fast streaming path plus a more accurate batch path under one model family - is the core engineering bet: rather than picking one operating point, Google shipped two tuned variants of the same underlying model.

Context-Aware Transcription: Screen and Chat History as a Signal

A capability easy to miss in the headline stats is that Gemini 3.5 Transcribe doesn't only listen - on Google Antigravity, it also looks. With user permission, the model pairs screen context and chat history to improve transcription accuracy for things like file names, agent thoughts, and active documents [5], effectively using visual and conversational context as a disambiguation signal the way a human transcriptionist would use surrounding context to guess an unfamiliar term. On the multi-speaker side, the model attributes speech with timestamps for up to three speakers in pre-recorded audio, with anything beyond three speakers still labeled experimental - a meaningful caveat for anyone planning to use it on group meetings or panel recordings rather than one-on-one calls [5]. The rollout strategy reinforces the surface-aware angle: the model already powers the Rambler dictation feature on Pixel 11-series phones and the Gemini app on macOS, with in-browser speech-to-text for any web form field coming to Chrome next [6].

The Verification Gap: Marketing Numbers vs Independent Scrutiny

There's a real tension between how this launch is being marketed and how it's holding up under outside inspection. Google's headline claims - the 70% latency improvement over Chirp 3 and the WER figures - trace back to measurements from Artificial Analysis that Google cites in its own launch materials; independent commentary flags that no outside party has yet run an adversarial head-to-head against open-weight transcription models [2]. Separately, other coverage notes that on raw accuracy, Gemini 3.5 Transcribe still trails specialized competitors: ElevenLabs' Scribe v2 and Microsoft's MAI-Transcribe-1.5 are both reported to post lower word error rates [1]. Community reception across platforms tracks that same split verdict rather than settling it. On X, the launch carried heavy official amplification - both Sundar Pichai and Google's own account promoted it directly - and at least one independent developer with early Google DeepMind access reported building and shipping a working app on the model, a sign of genuine builder enthusiasm alongside the official push. On YouTube, the most-watched and most authoritative video is Google's own developer walkthrough, framing the launch squarely as an API-first release aimed at developers building voice applications; third-party YouTube coverage exists too but is still same-day fresh. On Reddit, sentiment runs more measured and curious than hyped - one Reddit thread found the model beat Whisper and GPT-transcribe on a difficult Finnish-language test, while a separate user's own side-by-side testing found AssemblyAI came out ahead, and commenters confirmed the model is already live in Gboard's Rambler dictation feature. The pattern across analyst commentary, official amplification, and hands-on community testing is consistent: Gemini 3.5 Transcribe's real edge looks like latency, API integration, and developer-first positioning rather than a clean, undisputed accuracy win.

Historical Context

2025-03-17
Google announced Chirp 3, the prior-generation multilingual ASR model line under the separate Cloud Speech product family, which Gemini 3.5 Transcribe now supersedes.
2025
Chirp 3 Transcription reached general availability, expanding to 24 fully supported languages plus dozens more in preview, with diarization and automatic language detection as headline features.
2026-08-26
Gemini 3.5 Transcribe entered public preview, positioned as a direct successor to Chirp 3 with improved word error rate and 70% lower latency.

Power Map

Key Players
Subject

Google launches Gemini 3.5 Transcribe speech-to-text model

GO

Google / Google DeepMind

Developer and publisher of Gemini 3.5 Transcribe, controlling the model, API pricing, and rollout across first-party products

AG

Agora

Developer platform partner that integrated the model at launch for real-time communication apps

IN

IntelliTek Health

Healthcare technology customer deploying the model for real-time clinical transcription across primary care and specialties

LI

Lingopal

Customer building speech recognition products on the model

VI

vivo

Device OEM customer using the model for multilingual transcription and translation on its hardware

SO

Software Mansion / Fishjam

Developer platform partner that independently benchmarked the model against Google's published figures

Fact Check

7 cited
  1. [1] Google DeepMind Unveils Gemini 3.5 Transcribe for Enhanced Intelligent Speech-to-Text Processing
  2. [2] Gemini 3.5 Transcribe: Intelligent Transcription
  3. [3] Introducing Gemini 3.5 Transcribe
  4. [4] 247wallst.com Gemini 3.5 Transcribe coverage
  5. [5] Google introduces Gemini 3.5 Transcribe
  6. [6] Google announces Gemini 3.5 Transcribe
  7. [7] Google launches Gemini 3.5 Transcribe with smarter speech-to-text and 85-language support

Source Articles

Top 5

THE SIGNAL.

Analysts

Praised the model's speed and accuracy specifically with emails and numeric data, a common weak point for transcription models.

Mason Adams, Developer Evangelist, Agora
Positive customer testimonial

Says the model changes real-time clinical transcription workflows across care settings.

Martyn Molnar, CEO, IntelliTek Health
Positive customer testimonial

States the company's own internal measurements matched the performance numbers Google published.

Maciej Rys, VP Engineering, Software Mansion / Fishjam
Positive, independent benchmark validation

Frames the model's 0.40-second streaming latency as the commercially decisive figure in voice-agent procurement, positioning Google as newly competitive against specialized transcription vendors.

Industry analysis, 247wallst.com
Competitive, cautiously positive

Notes Google's 70% latency-improvement claim over Chirp 3 relies on measurements Google itself commissioned and cites, and that no outside party has run an adversarial benchmark against open-weight alternatives.

Industry analysis, orcarouter.ai
Mixed, skeptical of unverified claims
The Crowd

Say hello to Gemini 3.5 Transcribe! - Build apps that understand user speech / intent, even w/ multiple speakers! - Auto-detection of 85+ languages out of the box - Custom vocab adaptation for specialized jargon... SGTM:) API available now in @GoogleAIStudio and Gemini

@@sundarpichai4181

We're introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet It turns audio into precise transcription in 85+ languages, removing filler words like "ums" and "ahs" while handling self-corrections and capturing your intent so you can get things done using

@@Google5451

Google just launched Gemini 3.5 Transcribe, its most precise speech-to-text model yet. I built this app using Gemini 3.5 Transcribe and have been playing with the model for quite some time. Thanks to the Google DeepMind team for the early access. It's already available in the

@@ai_for_success163

Introducing Gemini 3.5 Transcribe

@u/Stoneonn213
Broadcast
How to build with Gemini 3.5 Transcribe

How to build with Gemini 3.5 Transcribe

Google's latest speech-to-text Gemini model offers a platter of new features

Google's latest speech-to-text Gemini model offers a platter of new features

Gemini 3.5 Transcribe: 2.6% WER at $0.005/Minute

Gemini 3.5 Transcribe: 2.6% WER at $0.005/Minute

Google launches Gemini 3.5 Transcribe speech-to-text model — AI News | Agentic Brew