Google Gemini 3.8 TTS: 30-second sample can replicate a voice.
动察 Beating AI News Flash: Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS, two text-to-speech models now available on the Gemini API and Google AI Studio. Flash focuses on audio quality, character performance, and long-text stability, while Lite leans toward high throughput, low latency, and low cost.
Both models can design new voices directly from text and can also replicate one's own voice or an authorized voice using about 30 seconds of reference audio. Users can control emotion, pacing, and accent sentence by sentence, and can even insert sounds such as laughter, sighs, and coughing. Two-person dialogue can also be generated directly from a single script. Flash supports 130 languages, and Lite supports 101.
Pricing is also lower than the previous generation. Until December 31, 2026, Flash audio output is $9 per million tokens, and Lite is $6; the previous generation Gemini 3.1 Flash TTS Preview was $20. Starting January 1, 2027, the standard prices for the two models will rise to $18 and $12.