Google has released two new text-to-speech models as part of its Gemini family. The company announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026.
The models turn written text into spoken audio. Google says they are its most expressive audio generation models so far.
They are available across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
Two Models Built for Different Jobs
Gemini 3.8 Flash TTS is designed for creative work and character design. Users can build new voices from scratch by describing them in plain language.
Google says it suits gaming, audiobooks, podcasts, and interactive media. Creators can direct each line with cues for acting, pacing, and dialect.
Gemini 3.8 Flash-Lite TTS is built for high-volume work at a lower cost. It targets dubbing, audio content creation, and voice agents.
The new models join other recent Gemini audio releases. These include 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.
Custom Voices and Safety Controls
Flash TTS can create voices in more than 100 languages and dialects. Users set the role, accent, and voice traits through text prompts.
The model also offers a library of more than 2,000 ready-made voices. That list includes regional varieties such as Mexican Spanish, Quebec French, and Scots English.
A voice replication feature can copy a voice from a 30-second audio sample. Before a voice is created, users must provide a verbal consent recording from the voice owner that matches the sample.
Every clip made by Gemini audio models carries a SynthID watermark. Google says this hidden mark helps people detect AI-generated speech.
Both models can produce hours of audio with little change in voice quality. They can also stage conversations between two speakers from a single script.
Scripts can include sounds like laughs, sighs, and gasps. Listening cues such as "mhm" and "yeah" can be added for timing.
A voice remixing tool is coming soon. It will let users adjust the pitch, pace, timbre, and accent of library voices.
Flash TTS took the top spot on Hume AI's Voice Design Benchmark with a score of 71.4. It also led in accent modeling with a score of 60.8.
On Hume AI's Overall Quality Index, Flash TTS ranked first and Flash-Lite TTS ranked second. Both models also placed near the top in blind tests on Voice Arena in languages including Japanese, Hindi, and Mexican Spanish.
Developer platforms Agora, LiveKit, Pipecat, and Vercel support the models through the Gemini API. Companies including Figma, HeyGen, Wondercraft, and Ollang are adding them for dubbing and voice agents.
Developers can now try the models in a new audio playground in Google AI Studio. It includes a voice design workspace and a two-speaker script editor.