Will it run?
Models

Google splits Gemini voice lineup into scale and reasoning tiers

By Rae Whitlock Clawpit staff

Google is dividing its voice-model portfolio into two distinct tracks: one built for scale and low cost, the other for multi-step reasoning that does not break the conversational flow. Announced yesterday, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are both available now through the Live API and the Gemini Enterprise Agent Platform, letting developers ship production-ready voice agents without wrestling with complex streaming infrastructure.

The reasoning model leads the benchmarks

Gemini 3.8 Live Extended Thinking takes the top spot on Artificial Analysis’s speech-to-speech quality index with a score of 82.6, and leads agentic task completion at 68.6 percent on τ-Voice and 35.1 percent on Sierra’s τ-Voice-banking benchmark. Its reasoning capability measures 97.7 percent on Big Bench Audio, all at what Google describes as a competitive price point against other frontier models. The technical differentiator is simultaneous thinking and speaking: the model emits early verbal cues such as “let me check that…” and streams a live progress narration while background tasks execute, preserving a natural feel even in long-running flows.

The base model targets scale and cost

Gemini 3.8 Live placed second in the Speech Agent Arena for user preference while remaining highly cost-efficient relative to its performance. It processes visual input in near real time, detects and switches automatically among 97 languages mid-conversation, and executes tool and API calls in the background without pausing the dialogue — the model acknowledges the request and keeps talking while the task finishes behind the scenes. That capability is critical for applications that demand low latency and a “live” feel without paying for a full reasoning cycle on every turn.

Developer ecosystem and enterprise partners

Platforms including Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents handle the heavy lifting of real-time streaming, freeing developers to focus on user experience instead of media plumbing. On the enterprise side, Salesforce, Genspark and Lumeris cite low latency, smooth flow and tool-calling as reasons for adoption. All audio emitted by the models carries an imperceptible SynthID watermark woven directly into the file to enable AI-content detection and reduce disinformation.

Immediate rollout to developers, enterprises and Search

Distribution began yesterday: for developers via the Gemini API and Google AI Studio; for enterprises in private preview on Gemini Enterprise, with Gemini Enterprise for Customer Experience coming soon; and for all users in Search Live. The Extended Thinking variant is also rolling out from yesterday, though Google did not provide exact regional timelines or final pricing beyond the general competitiveness claim.