Will it run?
Models

Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS with prompt-based voice design

By Rae Whitlock Clawpit staff
Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS with prompt-based voice design

Google yesterday released two new text-to-speech models in the Gemini Audio family: Gemini 3.8 Flash TTS, built for deep creative direction and character design, and Gemini 3.8 Flash-Lite TTS, aimed at high-volume, low-cost workloads. Both are available now through the Gemini API and Google AI Studio, API-only — no open weights for self-hosting — with enterprise access via Gemini Enterprise marked "coming soon." The model IDs in AI Studio's playground are gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts.

Until now Google offered only 30 preset voices. The new version moves to a far larger generative system: Flash TTS can create new voices from a prompt describing role, accent and vocal characteristics, and it works across more than 100 languages and dialects. Company demos include a Melbourne DJ, a monotone robot and a Japanese dragon. Developers also get a library of more than 2,000 production-ready voices with regional coverage such as Mexican Spanish, Quebec French and Scottish English. Custom voices can be saved and reused with minimal drift across projects, and a remix capability — adjusting timbre, pitch, pace and accent of a library voice via prompt — is marked "coming soon."

Both models accept directing instructions written directly inside the script, and Gemini can also steer delivery from natural cues in the text. Vocal quality, pacing and timbre hold over hours of continuous audio. There is native two-speaker support: a single script drives a multi-turn conversation with distinct, separated voices. Non-verbal vocal bursts such as <laughs>, <sigh> and <gasp> add conversational texture, and backchanneling — active-listening cues like |mhm| and |yeah| — controls response timing and comedy beats.

Voice cloning builds a consistent vocal profile from a 30-second sample. The sample must be the speaker's own voice or a voice the requester has rights to, and the process requires recording a verbal consent from the voice owner, matched against the original speaker. Every segment produced by the Gemini Audio models carries a SynthID watermark, an imperceptible stamp embedded directly in the audio. Cloned voices also carry C2PA content credentials. Google points to the Gemini 3.8 Audio model card for its broader safety approach.

According to Google's published data, Flash TTS ranks first overall on the Hume AI voice-design benchmark with a score of 71.4, and leads accent modeling at 60.8. On Hume's overall quality index, Flash TTS is first and Flash-Lite TTS second. In the blind-preference Voice Arena, the two models take the top spots in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi.