explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • Two models, two jobs
  • What voice design means in practice
  • Voice replication: capability and guardrails
  • Benchmarks: what Google claims
  • Price: the missing table
  • Where it ships
  • What people are asking
  • How to test it this week
  • Where this fits in the voice race
  • Use cases where the new features change the work
  • Implementation notes for developers
  • Bottom line
  • Related reading
← Back to blog

explainx / blog

Gemini 3.8 Flash TTS and Flash-Lite TTS: Voice Design, 2,000+ Voices and 30-Second Voice Cloning

Gemini, Google DeepMind, Text to Speech, Voice AI, Gemini API

Google's Gemini 3.8 Flash TTS and Flash-Lite TTS add voice design, 2,000+ voices, 100+ languages and consent-gated voice cloning. Details and regional limits.

Sep 24, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
Gemini 3.8 Flash TTS and Flash-Lite TTS: Voice Design, 2,000+ Voices and 30-Second Voice Cloning

Google is trying to make voice a design problem, not a menu. On September 23, 2026, Logan Kilpatrick announced Gemini 3.8 Flash and Flash-Lite TTS as "our new SOTA text to speech model," and Google's blog lays out what is inside: a new voice design experience, 2,000+ production-ready voices, voice replication, support for 100 languages, and a first-place claim on Hume AI's benchmarks.

This post summarizes what is confirmed, what to watch (cloning restrictions, pricing detail, benchmark claims), and how the release fits next to the rest of Google's audio push. For the model family context, read our Gemini 3.8 Flash launch and coding benchmarks and Gemini 3.8 Live.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
What launched?Gemini 3.8 Flash TTS (creative direction) and Flash-Lite TTS (scale, cost)
Voice design?Describe a voice in plain language; save it for reuse
Voices?2,000+ production-ready, incl. regional varieties (Quebec French, Scots English)
Languages?100+ languages and dialects
Cloning?From a 30-second sample plus spoken consent
Cloning blocked in?Illinois, Texas, EEA, UK, Switzerland, India
Watermark?SynthID on all audio
Coming soon?Voice remixing (timbre, pitch, pace, accent)
Cost?"Lower than 3.1 Flash TTS" per Logan Kilpatrick; no rate table in the announcement
Where?Gemini API, Google AI Studio; Flash-Lite in Google Vids; Flash in Gemini Notebook
Benchmarks?#1 on Hume AI Voice Design Benchmark (71.4), per Google

Two models, two jobs

Google split the release by use case.

Gemini 3.8 Flash TTS is for creative work: characters, narration, ads, games, audiobooks. It gets the deepest direction controls.

Gemini 3.8 Flash-Lite TTS is for volume: IVR, notifications, bulk content, apps where cost per minute drives the design.

This mirrors how Google positioned text models in the same generation, with Flash for capable-and-fast and Flash-Lite for cheap-and-fast. If you already route coding agents by model tier, see Claude Fable 5.1 and Gemini 3.8 Flash coding-agent routing. The same logic applies to voice: use Flash where quality matters, Flash-Lite where scale does.

What voice design means in practice

The headline feature is voice design. Instead of picking "Voice 14," you describe a character: role, accent, vocal traits. Google's blog says developers can describe a character in more than 100 languages and dialects, then save the resulting voice for reuse. That last part matters: a designed voice is only useful if it stays consistent across thousands of lines.

Related direction features in the release:

  • Line-by-line performance direction using natural script cues.
  • Dual-speaker scene staging, so a conversation can be generated as a scene instead of stitched clips.
  • Scripted vocal bursts and backchanneling: laughs, sighs, gasps, "mm-hm" interjections.
  • Long-form stability: hours of consistent audio without drift.

Example prompt style for a designed voice: "A warm, unhurried Scottish grandmother, mid-sixties, slightly amused, speaks in short sentences and pauses before punchlines." Then direct each line: "quietly, as if sharing a secret."

Voice replication: capability and guardrails

Voice replication uses a 30-second sample and matching verbal consent from the voice owner. Google says the system has "built-in consent verification," and that all generated audio includes SynthID, an imperceptible watermark for detecting AI-generated speech.

The regional exclusions are significant. Voice replication is unavailable in Illinois, Texas, the EEA, the UK, Switzerland and India. Those jurisdictions have strict biometric, publicity or AI-content rules, and Google is choosing to withhold cloning rather than risk non-compliance. If you build for global users, design a fallback: designed voices or stock voices where cloning is blocked.

Practical implications:

  1. Get written consent anyway. Spoken consent satisfies the tool, not necessarily your legal requirements.
  2. Store consent records alongside the voice ID.
  3. Disclose synthetic audio where required, and do not remove SynthID.
  4. Plan for region gating in your product, not just in your terms.

Benchmarks: what Google claims

Google says the models take the #1 overall spot on Hume AI's Voice Design Benchmark (71.4) and #1 and #2 on Hume AI's Overall Quality Index, plus top positions on Voice Arena in multiple languages. The announcement credits Hume's work; reporting notes Hume founder Alan Cowen was credited.

Treat these as vendor-reported. Voice quality is subjective and language-dependent, and rankings can change quickly. For an example of how a different provider was evaluated on an independent leaderboard, see Cartesia Sonic 3.6 on Artificial Analysis. The right approach is to test with your own scripts in your own languages.

Price: the missing table

Logan Kilpatrick said the new experience in Google AI Studio comes "with a lower cost than our previous 3.1 Flash TTS model." Google's announcement as fetched does not include a rate card or percentage. Until pricing is published on the Gemini API pricing page, do not assume a number. For cross-vendor cost work, see how we compare pricing in Gemini 3.7 Flash vs Grok 4.6 vs Sonnet 5 vs GPT-5.6.

Where it ships

  • Gemini API and Google AI Studio, for developers, starting now.
  • Google Vids, using Flash-Lite.
  • Gemini Notebook, using Flash.
  • Gemini Enterprise, coming soon through the API.

Google also announced a "Gemini Audio | At Night" event in San Francisco the day after launch, and Kilpatrick highlighted progress across TTS, Live, Lyria, Translate and Transcribe. Related earlier launches: Gemini 3.5 Transcribe for speech-to-text and Gemini 3.8 Live for real-time conversation.

What people are asking

"Another Flash? No more Pros?" One reply asked exactly that. Flash-tier models are where most volume and TTS production happen; a Pro-tier release is not part of this announcement.

"Is it better than ElevenLabs or Fish Audio?" No independent head-to-head is published yet. Fish Audio launched a competing controllable model the same day; see Fish Audio Drama 3 preview.

"Can I use it for Indian languages?" The announcement claims 100+ languages and dialects. Voice replication is blocked in India, but designed and stock voices may still be usable. For Indic-specific options, see Sarvam Bulbul v4 and Rumik OSS-1.

"Does it have an open version?" No. For local TTS, see Kokoro.

How to test it this week

  1. Create one designed voice for a real character and generate 20 lines. Check consistency.
  2. Try dual-speaker scenes with interruptions and backchanneling.
  3. Run your hardest text: numbers, names, mixed languages, code-switching.
  4. Compare Flash vs Flash-Lite on the same script; measure quality gap versus cost.
  5. Check SynthID handling in your pipeline. Make sure transcoding does not strip your compliance metadata.
  6. Read the region matrix before promising cloning to users.

Where this fits in the voice race

Three launches in 48 hours show the direction: OpenAI made ChatGPT Voice able to use plugins and GPT-6 (details), Fish Audio pushed fine-grained direction, and Google made voice design and cloning mainstream in its API. Meta's Muse also added voice and video at Connect (roundup). Speech is becoming the default interface for agents, and expressive output is becoming a commodity feature.

Use cases where the new features change the work

Games and interactive media. Voice design plus reusable saved voices means a studio can define a cast once, then generate thousands of lines that stay in character. Backchanneling and non-verbal cues (laughs, gasps) are the difference between a reading and a performance.

Podcasts and audiobooks. Long-form stability is the quiet requirement. Hours of audio that drift in tone or accent are useless; Google says the models keep quality consistent across long generations. Dual-speaker scene staging is a natural fit for interview-style or dialogue-heavy formats.

Localization. With 100+ languages and regional varieties (Quebec French and Scots English are examples Google gives), teams can match a market's accent instead of defaulting to a generic one. Test carefully: language count is a claim, and quality varies by language.

Customer-facing voice. Flash-Lite is aimed at high-volume use where cost matters. Pair it with a designed brand voice, and reserve Flash for premium content.

Education. Course creators can produce narrated lessons in multiple languages from one script. If you teach with video, a designed instructor voice makes updates easy: change one line, regenerate one line.

Implementation notes for developers

  • Store voice IDs and prompts. A designed voice is an asset. Save the prompt and the resulting voice reference in version control alongside your scripts.
  • Separate direction from text. Keep performance cues (how to say it) in a distinct field from the words (what to say). That makes A/B testing delivery simple.
  • Budget for retries. Even good TTS occasionally mispronounces names. Build a pronunciation review step for brand terms.
  • Keep SynthID intact. If your pipeline transcodes audio, check that you are not stripping metadata or processing in ways that defeat detection.
  • Gate cloning by region in code. Do not rely on users to know where cloning is unavailable.
  • Measure latency and cost per minute for Flash and Flash-Lite on your own text, because announcements rarely include both.

Bottom line

Gemini 3.8 Flash TTS and Flash-Lite TTS are the most complete voice toolkit Google has shipped: design, 2,000+ voices, cloning, scenes and watermarking in one API. The open questions are price, independent quality comparisons and the regional limits on cloning. If you build voice products, run a bake-off now.

Details reflect Google's September 23, 2026 announcement and press coverage. Availability, pricing and regional restrictions may change.

Related reading

  • Fish Audio Drama 3 preview
  • ChatGPT Voice gets plugins, GPT-6 and Work
  • Gemini 3.8 Flash launch and coding benchmarks
  • Gemini 3.8 Live and extended thinking
  • Gemini 3.5 Transcribe
  • Cartesia Sonic 3.6 TTS
  • Meta Connect 2026 roundup
  • Official: Google's Gemini 3.8 TTS announcement
Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 16, 2026

Google Launches Gemini 3.8 Live and 3.8 Live Extended Thinking

Gemini 3.8 Live and 3.8 Live Extended Thinking are Google's newest voice models, built for near real-time dialogue with background tool calls and live progress narration. One leads a speech-quality index outright; the other trades some of that fluency for deeper multi-step reasoning.

Sep 24, 2026

Fish Audio Drama 3 Preview: Directing AI Voice in Plain Language Instead of Audio Tags

Fish Audio introduced Drama 3, a preview text-to-speech model it calls the most controllable ever: describe tone and character in simple language, change voice mid-sentence, render multi-character scenes and regenerate a single word. Access is gated, pricing is unpublished. Here is what is confirmed, what to test, and how it compares.

Sep 8, 2026

Rumik-OSS-1: The First Open-Source Indic TTS With Emotion

Rumik Intelligence released Rumik-OSS-1 on September 8, 2026 — a 3-billion parameter open-weight TTS model built on Cohere's Tiny Aya Fire, claiming to be the first open-source Indic speech model with expressive emotion and parametric voice control across 22 languages.