Along with Gemini 3.8 Flash After the launch, unsurprisingly, Google also began updating its various AI models one after another. Earlier, it officially unveiled the next-generation text-to-speech models Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, only about 5 months after the previous generation, Gemini 3.1 Flash TTS, was released. In addition to further improving the naturalness and controllability of text-to-speech, this time it also added several new features, such as being able to clone a specified voice from only about 30 seconds of recording. It scored 71.4 in Hume AI’s Voice Design speech design benchmark, currently ranking first.

Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS: design a brand-new voice with a single sentence, and clone a human voice in just 30 seconds
Earlier, Google announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on its official blog. This upgrade can be divided into two main directions: first, “creating voices”—you no longer have to choose only from preset voices, but can design, clone, and save your own; second, “directing voices”—you can give instructions like a director for what tone and rhythm each sentence should be read in.
The “Making Sounds” section mainly consists of voice design and voice cloning, with the voice cloning feature supported by both Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.
You can describe characters, accents, and voice characteristics in natural language, and the model will generate an entirely new voice from scratch, supporting over 100 languages and dialects. If you don’t want to design your own, you can also use the ready-made voice library directly. This time, the voice library has expanded from the original 30 to over 2,000, and also includes regional accents such as Mexican Spanish, Quebec French, and Scottish English.
You can also “clone voices.” All you need to do is provide a 30-second recording, and it can clone a stable voice—whether it’s your own voice or a voice you have the right to use. To prevent it from being used to impersonate someone else’s voice, Google has added a consent verification step: users must separately upload a recording of the voice owner giving verbal consent. Only after the system confirms that this recording is of the same person as the reference recording will it create the voice.
Designed or cloned voices can be saved for management. According to the Gemini API developer documentation, up to 200 of these custom voices can be stored per project, with a retention period of 1 year.
In addition, Google also mentioned a “coming soon” voice-mixing feature: you can pick a voice from a voice library, then use prompts to fine-tune its timbre, pitch, speaking rate, and accent—for example, “add a bit of a Southern American accent” or “speak more gently.”
“Command Voice” is supported by both Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.
You can write stage directions for every line, or let Gemini decide how to deliver it based on the script, from a calm customer service tone to a hushed suspense scene. 3.1 relied on over 200 audio tags; this time the approach is closer to “directly writing a performance direction prompt,” which gives more natural control.
Google also mentioned that the new model can continuously generate several hours of audio, and that audio quality, pacing, and character voices remain consistent, with speaker voice drift rarely occurring, making it well suited for podcasts and audiobooks. Anyone who has worked on AI audiobooks knows that sometimes, by the latter half, the same character sounds like a different person.
It also natively supports two-person conversations. With a single script, you can create multi-turn dialogue between two people, and the two voices are clearly separated. The pacing of speaker changes is also very natural.
In terms of test data, the Hume AI voice design leaderboard is compared with ElevenLabs Voice Design v3 and Inworld Voice Design.
On the overall English score, Gemini 3.8 Flash TTS got 71.4, ElevenLabs got 70.8, and Inworld got 69.8, so the three were actually very close. The biggest difference was in accent: Gemini scored 60.8, while ElevenLabs scored only 45.4 and Inworld 35.8. In other words, when asked to produce specified accents like a “Melbourne accent” or a “Scottish accent,” Gemini is clearly more accurate. Gemini also held a slight lead in the multilingual category (3.82 vs. 3.65):

In Hume AI’s quality evaluation, Gemini 3.8 Flash TTS scored 0.920 overall, Flash-Lite TTS scored 0.914, and the previous generation Gemini 3.1 Flash TTS scored only 0.783, a clear improvement. It also leads all competitors. Gemini 3.8 Flash TTS is also first in multi-speaker dialogue and style tag control. The only category it lost was “Human-like variation,” which mainly assesses whether the variation between different speech outputs is close to the natural variation in how real people speak; ElevenLabs v3 received a perfect score of 5.00, while Gemini scored 4.58:

In Voice Arena’s human blind tests, Japanese stood out the most: Gemini 3.8 Flash TTS scored 1,232, while the previous generation 3.1 scored 1,148 and ElevenLabs v3 scored 1,048. Arabic also scored 1,204. More interestingly, for English, Brazilian Portuguese, and Vietnamese, the cheaper Flash-Lite actually scored the highest:

For security, in addition to the consent verification mentioned earlier, every piece of audio generated by the Gemini Audio model is embedded with Google’s SynthID watermark. This watermark is inaudible to the human ear, but it can be used to identify whether the audio was AI-generated. The voice cloning feature also supports C2PA content credentials.
Gemini 3.8 Flash TTS is rolling out starting today: developers can use it in the Gemini API and Google AI Studio, and general users can also experience it in Gemini Notebook.
As for API pricing, both models will have discounted pricing until December 31, 2026; starting January 1, 2027, text input, audio output, and batch processing prices will return to their original prices:
| Model (per million tokens) | Text input | Audio output | Batch processing output | Audio output starting in 2027 |
|---|---|---|---|---|
| Gemini 3.8 Flash TTS | $0.5 | 9 dollars | 4.5 US dollars | 18 US dollars |
| Gemini 3.8 Flash-Lite TTS | 0.5 US dollars | 6 US dollars | 3 US dollars | 12 US dollars |
| Gemini 3.1 Flash TTS (Preview) | 1 US dollar | 20 US dollars | 10 US dollars | — |
Source: KOCPC Chinese