Gemini 3.8 text-to-speech says hello
Sep 23, 2026, 8:25 AM · Google DeepMind

Google DeepMind ships Gemini 3.8 Flash and Flash-Lite TTS — generative voice design, replication with consent checks, and SynthID on every clip.
Why it matters
On September 23, 2026, Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS across AI Studio, the Gemini API, and related products, with Enterprise and Vids rolling out on staggered timelines. Creators can prompt new voices from scratch, pull from a 2,000-plus library, or replicate from a 30-second sample with consent verification.
Expressive speech at this level is useful AI for accessibility, localization, and interactive media. It is also the same stack that makes impersonation cheaper. Google is pairing capability with SynthID watermarking, C2PA credentials, and regional limits on replication — the right instincts, still unproven against determined misuse.
From the desk
We’re grading the launch on craft and containment. Flash TTS is aimed at character design and line-by-line direction — accents, pacing, dialect shifts, backchanneling, native two-speaker scenes, long-form with less speaker drift. Flash-Lite targets high-volume dubbing and voice agents. That split is how you serve both studio work and cost-sensitive production without forcing one model to do both poorly.
Google cites #1 overall on Hume AI’s Voice Design Benchmark (71.4) and accent modeling (60.8), plus top Voice Arena preference positions in languages including Japanese, Brazilian Portuguese, Vietnamese, MSA, Mexican Spanish, and Hindi, with support claimed for over 100 languages. Benchmarks are marketing until independent recreations land; still, the multilingual push is the part that earns our benefit of the doubt for useful AI — regional dubbing and agents that don’t sound like Mid-Atlantic defaults.
Voice replication is the sharp edge. A 30-second sample plus a matching verbal consent recording is a meaningful gate, and watermarking every Gemini Audio clip with SynthID plus C2PA is table stakes we want to see normalized. Replication is unavailable in Illinois, Texas, the EEA, UK, Switzerland, and India — a map of where law or caution is ahead of the product. That patchwork will frustrate developers and still won’t stop cross-border abuse if verification can be spoofed.
I’m watching whether consent verification holds under adversarial prompts, whether SynthID detection stays reliable after compression and remixing, and whether “coming soon” voice remixing from the library becomes the path around scarce talent. Partners named — Figma, HeyGen, Linguana, Wondercraft, and others — show where this lands first: dubbing, agents, and creator tools.
Useful voice AI should expand who can ship accessible audio. It should not quietly replace voice actors without consent or flood platforms with synthetic personas. Google shipped both the studio and the safety language. Now we need evidence the safety layer survives contact with the open web.
Context
DeepMind blog by Leland Rechis and Alan Cowen (Gemini Audio Team), September 23, 2026. Models follow earlier Gemini Audio releases including 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking. Flash TTS rolling out to API, AI Studio, Gemini Enterprise, Gemini Notebook, and Google Vids; Flash-Lite to API/AI Studio now, Enterprise soon, Vids for everyone.
Who feels it
- Creators and studios
- Generative design, dual-speaker staging, and long-form control cut production time for podcasts, games, and audiobooks — style consistency and rights clearance become the bottleneck.
- Developers / platforms
- API access via AI Studio and partners like Agora, LiveKit, Pipecat, and Vercel enables agent and dubbing pipelines; regional replication bans complicate global apps.
- Voice talent
- Consent verification and watermarking are protections on paper; contracts and detection tooling will decide whether replication stays licensed work or unpaid mimicry.
- Platforms and regulators
- SynthID + C2PA on every clip is a disclosure pattern to pressure peers; misuse cases will test whether watermarks survive real distribution.
What to watch
- Independent recreations of Hume / Voice Arena claims and long-form drift in production.
- Whether consent verification resists spoofed or coerced verbal consent samples.
- SynthID detectability after social-platform compression and edits.
- Enterprise rollout timing and how remixing features are gated when they ship.
Companies: Google