Models Verified
Gemini 3.8 Flash TTS and Flash-Lite TTS: voice design from prompts, 2,000+ voices
Google released two text-to-speech models billed as its most expressive audio generation models yet, moving voice creation from a fixed menu to a creative surface. Gemini 3.8 Flash TTS targets deep creative direction and character design — creating voices from scratch with natural-language prompts and directing every performance line by line with granular control over acting cues, pacing, dialect shifts and backchanneling — while Gemini 3.8 Flash-Lite TTS targets high-volume dubbing, audio content and production voice agents with fine-grained control over tone and pacing. Both offer 2,000+ production voices, 100+ languages and dialects including Mexican Spanish, Quebec French and Scots English, and voice replication from 30-second samples gated by consent verification, with built-in watermarking. Flash-Lite TTS replaces the gemini-3.1-flash-tts-preview model, and the migration guide requires moving turn-level directions into speech_metadata because 3.8 TTS now treats input text strictly as a verbatim transcript. Models are live in Google AI Studio, the Gemini API, Gemini Notebook and Google Vids, with Gemini Enterprise coming soon; launch partners include Figma, HeyGen, Linguana, Wondercraft, 99.co and Ollang, and developer platforms Agora, LiveKit and Pipecat.
GoogleGoogle DeepMind
Why it matters
Prompt-based voice design plus consent-gated 30-second cloning turns TTS into a creative tool, and pins the Flash-Lite tier as Google’s cost floor for production voice at scale.
View sources & tags2 sources · 7 tags
#gemini-3-8#tts#google#voice-cloning#audio#flash-lite#multilingual