Google is giving Gemini a much bigger voice. The company has launched two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, designed to make AI-generated speech more expressive, customizable, and useful for everything from audiobooks and podcasts to voice agents and video dubbing.
Google says the two models move text-to-speech beyond the usual collection of fixed voice presets. Instead, creators can describe the kind of voice they want using natural language, then direct how that voice performs individual lines.
Gemini 3.8 Flash TTS is the more creative of the two. It can generate entirely new voices based on prompts describing characteristics such as accent, role, and vocal style. Google says it supports more than 100 languages and dialects, allowing developers to create everything from fictional character voices to narrators with specific regional accents.
The model also has access to more than 2,000 production-ready voices. Google is additionally introducing voice replication, which can recreate a vocal profile from a 30-second audio sample, provided the user has the necessary rights to that voice. The company says the process includes consent verification, SynthID watermarking, and C2PA credentials.
That could make the technology particularly interesting for games, audiobooks and other projects that need consistent character voices. Creators can save custom voices and reuse them across projects rather than starting from scratch each time.
But creating the voice is only half of what Google is focusing on. Gemini 3.8 Flash TTS can also be directed line by line, letting creators specify how a particular sentence should sound. Prompts can control things such as pacing, emotion, acting style, dialect changes, and conversational reactions.
Google also supports nonverbal vocal cues and backchanneling, including things such as laughs, sighs, gasps, and short reactions such as “mhm” or “yeah.” That gives developers more control over the little details that make generated conversations sound less like someone reading text from a script.
The model can also generate long-form audio while maintaining voice quality and character consistency, according to Google. Native two-speaker scene staging allows creators to write dialogue for two characters in a single script while keeping their voices separate and handling conversational turn-taking.
Then there is Gemini 3.8 Flash-Lite TTS. Rather than focusing primarily on deep voice design, this version is built for high-volume and cost-efficient generation. Google says it is optimized for applications such as large-scale dubbing, audio content creation, and expressive voice agents, while still providing controls for tone, pacing, and other expressive details.
Google is positioning the models as part of its expanding Gemini Audio family, which already includes speech and real-time audio capabilities. The company says the new models performed strongly in its cited evaluations, including Hume AI’s Voice Design Benchmark and Voice Arena’s blind human preference evaluations. Google says Gemini 3.8 Flash TTS ranked first overall in Hume’s Voice Design Benchmark with a score of 71.4 and also led its accent-modeling evaluation with a score of 60.8.
The company says the models also performed strongly across languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi.
Safety is another major part of Google’s pitch. All audio generated by its Gemini Audio models receives an imperceptible SynthID watermark, which is intended to make AI-generated speech detectable. For voice replication specifically, Google says users must provide a verbal consent recording from the voice owner before a voice can be created.
Developers can start experimenting with Gemini 3.8 Flash TTS through the Gemini API and Google AI Studio. The model is also rolling out to Gemini Notebook, while Gemini 3.8 Flash-Lite TTS is available through the Gemini API, Google AI Studio, and Google Vids. Gemini Enterprise access is coming through its API.
There is one notable limitation: Google’s voice replication feature in AI Studio is not available in Illinois, Texas, the European Economic Area, the UK, Switzerland, or India.
The bigger shift here is that Google’s TTS technology is becoming less about choosing a voice and more about directing a performance. Instead of picking a preset narrator and hoping it fits the script, developers can describe a character, create a vocal identity, and then tell the model how that character should deliver every line.
For AI-generated audio, that is a meaningful change. And with Google making the tools available through AI Studio and the Gemini API, these capabilities are aimed at developers as much as they are at Google’s own products.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
