Google Unveils Record-Breaking Speech Generation Models

Google LLC released two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, on its cloud platform, according to company announcements. The algorithms feature highly similar application programming interfaces, designed to make side-by-side deployment simpler for developers building audio products.

Comparing Gemini 3.8 Flash TTS and Flash-Lite TTS Features

The two models serve different technical priorities within Google’s broader audio processing lineup, which already includes algorithms optimized for voice agents, transcription, and translation. According to Google, Flash-Lite TTS focuses on cost efficiency and inference speed. Flash TTS prioritizes higher audio quality at an increased price point, targeting applications such as audiobook creation.

Language support also separates the two algorithms at launch. Flash TTS generates speech across 130 languages, while Flash-Lite TTS supports 101 languages. Both models draw from a shared library of more than 2,000 prepackaged voices.

Pro Tip: Developers can test both models side-by-side with minimal code changes due to their nearly identical API structures.

Voice Customization and Synthetic Replica Controls

Developers can tailor the AI speakers using natural language prompts to adjust parameters like vocal timbre, accent, and pacing.

Google requires developers to secure explicit consent from the speaker before generating any replica. The company also announced plans for a third customization option in future updates, which will allow users to modify existing prepackaged voices.

To control delivery, both models accept oratory cues added to individual lines of a script. These cues trigger specific audio elements during playback, including non-lexical vocalizations and natural pacing shifts.

Benchmark Performance and Industry Evaluation

Google evaluated the models using an audio quality benchmark developed by startup Hume AI Inc., where Flash TTS secured the first spot and Flash-Lite TTS earned second place. Additionally, the models outperformed competing algorithms across multiple language-specific versions of Voice Arena, a benchmark measuring output quality through human feedback.

“These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids,” Google staffers Leland Rechis and Alan Cowen wrote in a blog post.

Did you know? Voice Arena relies entirely on human evaluators to score text-to-speech output, providing a real-world measure of naturalness compared to purely automated metrics.

Frequently Asked Questions

What is the main difference between Gemini 3.8 Flash TTS and Flash-Lite TTS?

Flash-Lite TTS is optimized for cost efficiency and inference speed across 101 languages, whereas Flash TTS offers higher audio quality, supports 130 languages, and is designed for premium tasks like audiobooks.

How many prepackaged voices are available in the new models?

Both models provide access to a library of more than 2,000 prepackaged voices.

Can developers create custom voices from audio samples?

Yes. Developers can generate a synthetic voice replica using a 30-second audio sample, provided they secure the speaker’s consent beforehand.

How do the models handle delivery and pacing?

Developers can insert oratory cues into script lines to generate non-lexical vocalizations and pacing shifts during speech generation.


Support open technology coverage by engaging with theCUBE community. Connect with more than 11,400 tech and business leaders in the Alumni Trust Network to share intelligence and discuss breakthroughs in AI and cloud infrastructure.

Leave a Comment