Our Blog

Blog Index 

Google Launches Gemini 3.8 Live and 3.8 Live Extended Thinking: Real-Time Voice AI That Reasons While It Speaks

Posted on 18th Sep 2026 06:04:13 in Artificial Intelligence, Machine Learning

Tagged as: Gemini 3.8 Live, Google DeepMind, Voice AI

Google's push to make artificial intelligence feel less like software and more like a conversation reached a new milestone this week. On September 15, 2026, Google DeepMind launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — a pair of voice-first models built for real-time dialogue in which reasoning, tool use and speech all happen at the same time instead of taking turns.

Both models began rolling out immediately: for developers in the Gemini API and Google AI Studio, for enterprises in private preview inside Gemini Enterprise, and for everyone through Search Live. Extended Thinking is also arriving in the Gemini app and across Google Workspace in Docs, Gmail and Keep. The launch is a step change from text-first assistants — Google calls these its "most advanced live dialogue models yet," and the independent benchmarks appear to agree.

Two Models, Two Jobs — Both Built Around Conversation

Gemini 3.8 Live is the workhorse. Built for scale and cost efficiency, it combines conversational intelligence with fluid dialogue and visual grounding. Google says it handles complex reasoning, real-time visual context and background task execution without interrupting the conversation. In a demonstration published with the launch, the model guides an employee through onboarding in real time, answering live questions using what it can see, and plays a game of chess against a person while reasoning out loud about the board in front of it.

Gemini 3.8 Live Extended Thinking targets high-complexity work. It reasons and speaks simultaneously — using early verbal cues such as "Let me check that…" to acknowledge a request naturally, then narrating its progress through multi-step background tasks while the conversation keeps flowing. In other demos it turns a raw sketch and spoken feedback into functional React components, coordinates restaurant reservation slots and table details asynchronously, and builds complete business plans through natural speech alone.

The bigger idea is that voice is no longer a thin layer on top of a text model. These are native speech-to-speech systems: the model listens, thinks and talks in one continuous stream, which is why it can acknowledge you instantly and finish the hard part in the background.

The Benchmarks Behind the Claim

Google backed the launch with an unusually detailed set of independent evaluations. On the Artificial Analysis Speech-to-Speech Quality Index, Gemini 3.8 Live Extended Thinking captured the number one overall position with a score of 82.6. The company also reported:

  • 68.6% on ?-Voice and 35.1% on Sierra's ?-Voice-banking benchmark — both measuring agentic task completion over live voice, where Extended Thinking claims the top spot.
  • 97.7% on Big Bench Audio, a benchmark of audio reasoning ability.
  • A position on the Pareto frontier of ServiceNow's EVA-Bench, which evaluates voice agents on balancing accuracy with conversational quality in complex workflows.
  • A second-place finish in the Speech Agent Arena for Gemini 3.8 Live, which Google frames as the cost-efficient option built for scale.

In plain terms: the model that talks fastest is not the one that thinks best, and vice versa — Google is shipping both, and letting builders pick based on the job at hand.

How It Works Under the Hood

The developer documentation fills in the technical picture. Both models accept text, images, audio and video as inputs and produce text and audio as output, with an input context of 131,072 tokens and an output limit of 65,536 tokens. Asynchronous function calling is now the default: the model fires off API and tool calls in the background while continuing to stream audio, so it can acknowledge a request and keep chatting while the task completes. That is the machinery behind the reservation-booking and onboarding demos.

Other capabilities Google highlights for builders include visual context (the agent understands what you say and what you show it), alphanumeric precision for reading back confirmation codes and claim numbers, automatic switching between 97 supported languages mid-conversation with accent consistency, and incremental content updates that merge real-time audio with structured data.

Extended Thinking goes further. It runs background reasoning while streaming continuous audio, exposes a configurable thinking level (low, medium or high), and uses an interaction status field so client apps know whether the server is still processing or idle. Its function calling is asynchronous-only — a hint at how strongly Google now assumes that "talking while working" is the default mode for voice agents. For developers already on the older live models, Google published migration notes covering the changes.

Pricing and Where You Can Try It Today

Cost has been the quiet blocker for production voice agents, and Google attacked it directly. Both models are priced at $0.005 per minute of audio input and $0.018 per minute of audio output — an estimate based on $3 per million input tokens and $12 per million output tokens — which Google describes as a competitive price point against other frontier voice models.

Developers can start in the Gemini API and Google AI Studio, or build on top of integration partners that handle media streaming infrastructure: Agora, Fishjam, LiveKit, LangChain, Pipecat, Vercel and Vision Agents. Enterprises get Gemini Enterprise in private preview, with Gemini Enterprise for Customer Experience coming soon; Google also named Salesforce, Genspark and Lumeris among companies it is partnering with around the new models. Alongside the launch, Google recapped its wider audio suite, including Gemini 3.5 Transcribe for 85-plus-language transcription, which it says achieves a 4.0% word error rate in streaming mode.

Two trust features round out the release. All audio generated by Google's AI products is watermarked with SynthID, an imperceptible marker woven into the output so AI-generated speech stays detectable, and a public model card documents the safety approach for the new audio models.

Why It Matters for Businesses and Developers

Voice has quietly become the most contested surface in AI, and this launch shows why. Every call centre, booking desk, onboarding flow and field-service app in India runs on human conversation — and speech-to-speech models finally make it practical to put an agent on that conversation without a visible lag chain of transcribe-then-generate-then-speak. The 97-language coverage and mid-sentence language switching matter enormously for a country where a customer may greet you in Hindi, switch to English and finish in Marathi.

Ambr AI, one of the launch partners, said in a testimonial that switching its enterprise negotiation-training simulations to Gemini 3.8 Live Extended Thinking made them "faster, more expressive and better at handling complex conversations," and let it deliver training in over 70 languages through a single integration. That is the shape of the opportunity: not gimmick voice assistants, but training, support and sales workflows rebuilt around live agents that can think while they talk.

The competitive pressure is now on pricing and reasoning-in-conversation, not just benchmark scores. With Extended Thinking at the top of the speech-to-speech quality index and the standard Live model built for scale, Google has drawn a line that the rest of the frontier labs — and the thousands of businesses watching their AI bills — will have to answer.

Sources

whatsapp me