Gemini 3.8 Flash TTS lets you build a voice from a 30-second sample
Gemini 3.8 text-to-speech says hello

Google's new Gemini 3.8 Flash TTS and Flash-Lite TTS models turn voice generation into a creative studio. You can design voices from scratch with natural-language prompts, replicate a voice from a 30-second sample, and direct line-by-line delivery with cues like <laughs> or |mhm|. Both models top Hume AI's quality benchmarks, support over 100 languages, and ship with SynthID watermarking and consent verification.
Scale up from 30 original voices to an infinite library.