Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS with custom voice design, voice replication, and line-by-line performance direction
Google DeepMind launched two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, available starting today in Google AI Studio (and via Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids). New capabilities include generative voice design from natural-language prompts, an expanded library of 2,000+ production-ready voices, voice replication from a 30-second sample with consent verification, line-by-line performance direction (pacing, emotion, dialect, backchanneling), native two-speaker scene staging, and SynthID watermarking plus C2PA credentials on generated audio. Voice remixing is announced as "coming soon" (not yet available).
Entities: Google DeepMind, Gemini 3.8 Flash TTS, Gemini 3.8 Flash-Lite TTS, Google AI Studio, Gemini API, Gemini Enterprise
1 primary
What happened
Google DeepMind launched two text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, live today in Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids. New features include prompt-based voice design, a library of over 2,000 preset voices, voice cloning from a 30-second sample (with consent verification), line-by-line direction of pacing, emotion and dialect, two-speaker scene staging, and SynthID/C2PA watermarking. No pricing has been disclosed, and a "voice remixing" feature is announced but not yet available.
Why it matters
This affects developers and enterprises building voice products, and consumer-facing tools that use Gemini's audio stack (Vids, Notebook). It gives teams already on Gemini more granular control over synthetic voice output without switching vendors, which matters for cost and integration decisions if pricing turns out competitive. The impact is bounded: TTS is a mature, crowded market with strong incumbents (ElevenLabs, OpenAI, Hume), and this is an iterative upgrade over Gemini 3.1 Flash TTS rather than a new capability category.
What is noise
The "static presets into a dynamic creative studio" framing is standard launch copy, not a real category shift, since prompt-based voice design and cloning already exist at competitors. The benchmark wins cited (Hume AI Voice Design Benchmark, Voice Arena) are vendor-selected comparisons with no independent verification, and "most expressive models yet" is a self-referential claim against Google's own prior model. Missing context: no pricing, no latency or cost-per-character figures, and the flagship "voice remixing" feature is not actually shipping yet.
Watch next
- 01Pricing and rate limits for the two models once published, compared to ElevenLabs and OpenAI TTS per-character costs
- 02Independent (non-Hume, non-Google) benchmark or blind listening test results within the next 4-8 weeks
- 03Actual ship date and feature scope of the 'coming soon' voice remixing capability
- 04Adoption signals: third-party apps or Gemini Enterprise customers publicly integrating the new voice replication or performance-direction features
Coverage
2 storiesMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- OpenAI launches GPT-6 Sol and Luna at half the token price of GPT-5.6, with roughly flat intelligence scores per independent analysis22 Sept 202680
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680