Google has introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two new text-to-speech models in the Gemini family. The models extend Google’s audio generation capabilities across its developer platforms and AI products, with availability varying by model and service.
The two models follow Gemini 3.5 Live Translate, Gemini 3.5 Transcribe, Gemini 3.8 Live, and Gemini 3.8 Live Extended Thinking as part of the Gemini Audio family.
Gemini 3.8 Flash TTS and Flash-Lite TTS
Gemini 3.8 Flash TTS focuses on voice creation and character design, while Gemini 3.8 Flash-Lite TTS targets high-volume speech generation. Flash-Lite is optimized for high-volume dubbing, audio content creation, and expressive voice agents.
With Gemini 3.8 Flash TTS, users can create voices from scratch using natural-language prompts and customize role, accent, and voice characteristics across more than 100 languages and dialects. The model provides more than 2,000 production-ready voices, including regional varieties such as Mexican Spanish, Quebec French, and Scots English.
Other voice capabilities include:
- 30 original voices: The voice library starts with 30 original voices and can be expanded with custom voices.
- Voice replication: A 30-second audio sample can be used to recreate a vocal profile of the user’s own voice or a voice they have the rights to use.
- Voice management: Custom voices can be saved and managed for continued use across projects.
- Voice remixing: Coming soon, users can select a library voice and adjust its timbre, pitch, pace, and accent through prompts.
Voice replication requires a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. Google also uses SynthID watermarking and C2PA credentials for voice replication.
Line-by-line performance control
After selecting a voice, both TTS models allow users to control individual lines through stage directions and natural script cues. These controls cover acting cues, pacing, tone, dialect shifts, expressive nuances, and backchanneling, while the models also support long-form generation with voice quality, natural pacing, and character timbre maintained across hours of continuous audio with minimal speaker drift.
Native two-speaker scene staging allows multi-turn conversations to be created from a single script while keeping the voices separate with conversational turn-taking. Scripts can also include non-verbal cues and active-listening responses, including:
- Non-verbal cues:
<laughs>,<sigh>, and<gasp> - Active-listening interjections:
|mhm|and|yeah|
Performance and language support
Google reports results for the two models across voice design, accent modeling, overall quality, and human preference evaluations. On Hume AI’s Voice Design Benchmark, Gemini 3.8 Flash TTS recorded an overall score of 71.4 and an accent modeling score of 60.8, securing the #1 overall position on the benchmark.
The models also secured the following positions on Hume AI’s Overall Quality Index:
- Gemini 3.8 Flash TTS: #1
- Gemini 3.8 Flash-Lite TTS: #2
Google says the models show improvements in long-form content and dual-speaker screenplay control compared with Gemini 3.1 Flash TTS. In blind human preference evaluations on Voice Arena, both models secured top positions among competitors in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish, and Hindi, with support for more than 100 languages.
Safety and transparency
Google says its voice creation and replication capabilities include safeguards for voice talent, identity, and content transparency. For voice replication, users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created, while every audio clip generated by Google’s Gemini Audio models carries a SynthID watermark embedded directly into the audio output.
Google says the watermark keeps AI-generated speech detectable and helps address misinformation. Further information about its safety and responsibility approach is available in the model card.
Google AI Studio
Google AI Studio includes an audio playground for the new speech generation capabilities. Developers can create vocal identities from scratch using prompts or replicate their own voice, then use the voices in a dual-speaker screenplay editor for line-by-line delivery control.
Voice replication through Google AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland, or India.
Developer and company integrations
The Gemini API supports speech generation integrations with developer platforms including Agora, LiveKit, Pipecat, and Vercel. Google also said Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang are integrating the new TTS models for applications including global dubbing, media localization with regional accents, and conversational voice agents.
Availability
Gemini 3.8 Flash TTS is rolling out starting today with the following availability:
- Developers: Gemini API and Google AI Studio
- Enterprises: Coming soon via API in Gemini Enterprise
- Everyone: Gemini Notebook
Gemini 3.8 Flash-Lite TTS is also rolling out starting today with the following availability:
- Developers: Gemini API and Google AI Studio
- Enterprises: Coming soon via API in Gemini Enterprise
- Everyone: Google Vids