Gemini 3.8 Live and Extended Thinking launch
TECH

Gemini 3.8 Live and Extended Thinking launch

29+
Signals

Strategic Overview

  • 01.
    Google DeepMind launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, calling them its most advanced live dialogue models yet.
  • 02.
    Live is built for scale and cost efficiency with fluid conversational dialogue and visual grounding, while Extended Thinking targets high-complexity tasks that need multi-step reasoning.
  • 03.
    Both models automatically detect and switch between 97 supported languages in the middle of a conversation.
  • 04.
    The models can execute tools and API calls in the background without interrupting the live conversation, acknowledging a request and continuing to talk while the task finishes.
  • 05.
    Extended Thinking reasons and speaks at the same time instead of pausing to think before responding.
  • 06.
    Both models are part of the Gemini 3 series, natively multimodal, accepting audio, images, video and text up to a 128K token context window.
  • 07.
    All AI-generated audio from the models is watermarked with SynthID, and they are available to developers via the Gemini API and Google AI Studio, in private preview for Gemini Enterprise, and to consumers through Search Live, Gemini Live, and Workspace apps including Docs, Gmail and Keep.

Live vs Extended Thinking: The Mechanics Behind Talking While Thinking

Gemini 3.8 Live and its Extended Thinking sibling are built on the same Gemini 3 multimodal backbone, accepting text, images, audio and video across a 128K-token input window [2]. What separates the two variants is how they handle latency versus depth: Live is tuned for scale and cost efficiency, pairing fluid dialogue with visual grounding, while Extended Thinking is built to reason and speak simultaneously rather than pausing to think before it answers [1]. Both models can kick off tools and API calls in the background without breaking the conversation - acknowledging a request, then continuing to chat while the task finishes off-mic [1]. Both models also auto-detect and switch between 97 languages mid-conversation [1]. Independent developer commentary places the underlying architecture in the same family as OpenAI's speech-to-speech GPT-Live models, reachable over a comparable WebSocket API rather than a bespoke integration [4].

The Voice AI Price War: Undercutting GPT-Live-1 Astra and Grok Voice

Google is not just claiming a quality win - it is pricing to match. Extended Thinking costs $3.50 per hour of audio, versus $5.83 per hour for OpenAI's GPT-Live-1 Astra and $4.80 per hour for xAI's Grok Voice Think Fast 2.0, while topping both rivals on Artificial Analysis' Speech to Speech Quality Index at 82.6 versus 81.5 and 81.3 [3]. The same pattern holds on agentic capability: Extended Thinking leads the tau-Voice benchmark at 68.6%, ahead of GPT-Live-1 Astra's 67.9% and Grok Voice's 56.5%, and leads Sierra's banking-specific tau-Voice-banking leaderboard at 35.1%, ahead of GPT-Live-1 Astra's 32.0% and xAI-Realtime's 16.5% [3]. Base-tier Gemini 3.8 Live is priced even lower, at $0.84 per hour of input audio, undercutting prior-generation options like GPT-Realtime 2.1 by a wide margin [5]. The combination of a top benchmark slot and the cheapest per-hour rate is a deliberate two-front strategy: win on quality metrics enterprises can point to, and win on unit economics for anyone building a high-volume voice agent.

The Reasoning-Speed Paradox: Why Some Users Prefer the Faster Model

Google's own marketing draws a clean line: Live is the fast, cheap option and Extended Thinking is the smarter one, with the implicit assumption that more reasoning is strictly better. The benchmark data mostly backs that framing - Extended Thinking's 82.6 Speech to Speech Quality Index score and its wins on tau-Voice and tau-Voice-banking put it ahead of rival models from OpenAI and xAI [3]. But the launch reception among actual users tells a messier story. Community discussion has been split on the live voice persona itself, with some testers calling it chipper and annoying and others finding it charming, and - more pointedly - one thread noted Extended Thinking scoring worse than base Live on certain interactive leaderboards, with the working theory being that live conversation rewards speed over deliberation and users don't want a voice agent that visibly pauses to think before replying. That tension - a model engineered to reason while it talks, running into users who prefer it not to reason much at all - is arguably the most interesting open question the launch leaves unresolved: whether smarter and better voice agent are actually the same axis.

From I/O Demo to Production: Project Astra's Path to Gemini 3.8 Live

Gemini 3.8 Live doesn't appear out of nowhere. Google first introduced a talk-to-Gemini mode at I/O 2024, paired with the viral Project Astra demo of near-real-time multimodal AI [6]. A year later, Astra's low-latency multimodal techniques were folded into Google Search, the Gemini app and developer tools, laying the technical groundwork that this launch builds on [7]. What's changed since then is scope: this release folds background tool execution, 97-language switching, and now a genuine choice between a cheap fast model and a high-reasoning one into a single production-ready API and AI Studio surface, alongside Workspace apps like Docs, Gmail and Keep [1]. The tradeoff for that expanded surface area is that Google's own model card acknowledges the models may still carry general foundation-model limitations such as hallucination, a risk that matters more in an always-on, tool-executing voice agent than in a chat window [2]. For businesses that already run Gemini text models for agentic workflows, having a same-family voice option priced well below rival speech models is the more immediate commercial hook [5].

Historical Context

2024-05-14
Google announced Gemini Live at I/O 2024, letting users talk to Gemini via the mobile app, alongside a viral Project Astra demo of near real-time multimodal AI.
2025-05-20
Project Astra's low-latency, multimodal capabilities were extended into Google Search, Gemini, and developer tools, laying groundwork for the real-time voice features now in Gemini 3.8 Live.
2026-09-15
Google DeepMind launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, its newest real-time voice models with background tool calling and 97-language support.

Power Map

Key Players
Subject

Gemini 3.8 Live and Extended Thinking launch

GO

Google DeepMind

Developer of Gemini 3.8 Live and Extended Thinking; announced the launch as its best conversational AI, extending the Gemini 3 model family into real-time voice.

OP

OpenAI (GPT-Live-1 Astra)

Direct competitor whose speech-to-speech model was outranked by Gemini 3.8 Live Extended Thinking on the Artificial Analysis Speech-to-Speech Quality Index (82.6 vs 81.5) and priced higher at $5.83/hour vs Google's $3.50/hour.

XA

xAI (Grok Voice / Think Fast 2.0)

Competitor whose Grok Voice Think Fast 2.0 (High) scored 81.3 on the same quality index and costs $4.80/hour, more expensive and lower-scoring than Gemini 3.8 Live Extended Thinking.

AR

Artificial Analysis

Independent benchmarking organization whose Speech to Speech Quality Index and tau-Voice agentic benchmark rankings underpin Google's third-party performance claims for the new models.

Fact Check

7 cited
  1. [1] Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking
  2. [2] Gemini 3.8 Audio Model Card
  3. [3] Google Releases Gemini 3.8 Live-Extended Conversational Model, Claims Better Performance Than Rivals At Lower Price
  4. [4] Gemini 3.8 Live and 3.8 Live Extended Thinking
  5. [5] Google Gemini 3.8 Live Extended Thinking Announcement
  6. [6] Gemini Live announced at Google I/O 2024
  7. [7] Project Astra comes to Google Search, Gemini and developers

Source Articles

Top 5

THE SIGNAL.

Analysts

Frames Gemini 3.8 Live as structurally comparable to OpenAI's GPT-Live family, noting it is reachable through a practical WebSocket implementation similar in shape to existing speech-to-speech offerings.

Simon Willison
Independent AI developer and commentator
The Crowd

Introducing our most advanced Gemini Audio models yet 🗣 Gemini 3.8 Live and 3.8 Live Extended Thinking let you speak, collaborate, and execute tasks seamlessly, meaning conversing with AI just got a lot more natural. So, what's the difference between these two models? Let's

@@GoogleAI2853

Google has released Gemini 3.8 Live, its new Speech to Speech model, with the Extended Thinking (High) variant debuting at #1 on the Artificial Analysis Speech to Speech Index at 82.6, and #1 on our Tau Voice benchmark implementation at 68.6% Gemini 3.8 Live is @GoogleDeepMind's

@@ArtificialAnlys1200

My mind is f*** blown. I built this live insurance claim agent that can see, talk, think and draw in real-time using the new Gemini 3.8 LIVE. Even switched my language to Hindi mid-call and it still worked. VOICE AI can't be more real. Made it 100% open-source.

@@Saboo_Shubham_1009

Gemini 3.8 Live & Gemini 3.8 Live Extended Thinking

@u/gibbonwalker268
Broadcast
What's new in the Gemini Live API

What's new in the Gemini Live API

Gemini 3.8 Live VE y HABLA en tiempo real (y da miedo)

Gemini 3.8 Live VE y HABLA en tiempo real (y da miedo)

【環境最強】リアルタイムAI『Gemini 3.8 Live』が最強コスパで登場!精度も体験も良いので解説

【環境最強】リアルタイムAI『Gemini 3.8 Live』が最強コスパで登場!精度も体験も良いので解説