Signum
Feed
Useful signal23 Sept 2026high confidence

Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS with custom voice design, voice replication, and line-by-line performance direction

Google DeepMind launched two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, available starting today in Google AI Studio (and via Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids). New capabilities include generative voice design from natural-language prompts, an expanded library of 2,000+ production-ready voices, voice replication from a 30-second sample with consent verification, line-by-line performance direction (pacing, emotion, dialect, backchanneling), native two-speaker scene staging, and SynthID watermarking plus C2PA credentials on generated audio. Voice remixing is announced as "coming soon" (not yet available).

CapabilityAccessEconomics

Entities: Google DeepMind, Gemini 3.8 Flash TTS, Gemini 3.8 Flash-Lite TTS, Google AI Studio, Gemini API, Gemini Enterprise

66Useful signal
2 sources
1 primary
Was this useful?
01

What happened

Google DeepMind launched two text-to-speech models, Gemini 3.8 Flash TTS and Flash-Lite TTS, live today in Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids. New features include prompt-based voice design, a library of over 2,000 preset voices, voice cloning from a 30-second sample (with consent verification), line-by-line direction of pacing, emotion and dialect, two-speaker scene staging, and SynthID/C2PA watermarking. No pricing has been disclosed, and a "voice remixing" feature is announced but not yet available.

02

Why it matters

This affects developers and enterprises building voice products, and consumer-facing tools that use Gemini's audio stack (Vids, Notebook). It gives teams already on Gemini more granular control over synthetic voice output without switching vendors, which matters for cost and integration decisions if pricing turns out competitive. The impact is bounded: TTS is a mature, crowded market with strong incumbents (ElevenLabs, OpenAI, Hume), and this is an iterative upgrade over Gemini 3.1 Flash TTS rather than a new capability category.

03

What is noise

The "static presets into a dynamic creative studio" framing is standard launch copy, not a real category shift, since prompt-based voice design and cloning already exist at competitors. The benchmark wins cited (Hume AI Voice Design Benchmark, Voice Arena) are vendor-selected comparisons with no independent verification, and "most expressive models yet" is a self-referential claim against Google's own prior model. Missing context: no pricing, no latency or cost-per-character figures, and the flagship "voice remixing" feature is not actually shipping yet.

04

Watch next

  1. 01Pricing and rate limits for the two models once published, compared to ElevenLabs and OpenAI TTS per-character costs
  2. 02Independent (non-Hume, non-Google) benchmark or blind listening test results within the next 4-8 weeks
  3. 03Actual ship date and feature scope of the 'coming soon' voice remixing capability
  4. 04Adoption signals: third-party apps or Gemini Enterprise customers publicly integrating the new voice replication or performance-direction features

Coverage

2 stories

More capability signals

Full feed →