explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • What is confirmed
  • What is not confirmed
  • Why "no audio tags" is a big deal
  • The three features, and how to test each
  • How it compares to other 2026 voice models
  • What people are asking
  • A practical evaluation checklist
  • When Drama 3 could matter most
  • Sample direction prompts to try once you have a key
  • Bottom line
  • Related reading
← Back to blog

explainx / blog

Fish Audio Drama 3 Preview: Directing AI Voice in Plain Language Instead of Audio Tags

Fish Audio, Text to Speech, Voice AI, Audio Models, Developer Tools

Fish Audio Drama 3 (preview) lets you direct tone, pacing and characters in plain language, shift voice mid-sentence and fix one word. Known facts and tests.

Sep 24, 2026·8 min read·Yash Thakker
add explainx.ai
go deep
Fish Audio Drama 3 Preview: Directing AI Voice in Plain Language Instead of Audio Tags

Text-to-speech has a control problem. The best models sound human, but getting the exact delivery you want usually means learning a private markup language: bracketed emotion tags, SSML, pause markers, stress annotations. Fish Audio's new Drama 3 preview is a bet that the markup should disappear.

On September 23, 2026, Fish Audio announced Drama 3 (preview) as "the most controllable TTS model ever," where you "describe the tone, pacing, and character in simple language, without thinking in audio tags." The announcement lists three headline abilities: shift your voice mid-sentence, generate a multi-character scene, and fix only a single word. This guide covers what is confirmed, what is not, and how to evaluate it before you build on it. For background on the company, see Fish Audio's $52M seed and S2.1 Pro.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
What is it?Drama 3 preview, a text-to-speech model from Fish Audio
Key ideaDirect delivery with plain-language descriptions, not audio tags
Headline featuresMid-sentence voice shifts, multi-character scenes, single-word regeneration
API namedrama-3-preview (per reporting)
AccessGated: comment "DRAMA" on the announcement to be sent a key by DM
PricingNot published
StatusPreview; expect changes
Best fitDialogue, games, audiobooks, ads and any workflow that edits takes

What is confirmed

The announcement is short, so the confirmed list is short too:

  • Drama 3 is a preview, positioned as the most controllable TTS model Fish Audio has shipped.
  • Control is done through simple natural language, describing tone, pacing and character.
  • It can change voice mid-sentence, producing a shift within one line rather than needing separate clips stitched together.
  • It can render a multi-character scene in a single pass.
  • It can fix a single word without regenerating the rest of a take.
  • API keys are being distributed through a comment-and-DM flow.

Reporting from AlphaSignal adds that the model is exposed as drama-3-preview next to the S2 family in Fish Audio's TTS API, and that keys are gated rather than self-serve. Fish Audio's own documentation lists multi-speaker synthesis for the S2 family and the Drama 3 preview model, which is consistent with the multi-character claim.

What is not confirmed

Be careful with anything beyond the list above. As of publication:

  • No public price. One commenter asked about pricing on the announcement and got only a "check your DM." Ask for rates in writing.
  • No published latency, streaming or language coverage for Drama 3.
  • No independent benchmark. "Most controllable" is a claim from the vendor; there is no third-party score yet.
  • No details on voice cloning, consent controls or watermarking for this model.
  • No terms of service or commercial-use notes for the preview.

Treat the demo video as marketing until you can generate your own clips.

Why "no audio tags" is a big deal

To see why the announcement matters, compare how control works today.

Tag-based control. S2-family Fish Audio models accept natural-language tags such as [whisper], [excited] or [angry] placed inside the text. That is already friendlier than SSML, and it enables sub-word control of prosody and emotion. But you still have to know which tags exist, where to place them, and how they interact.

Description-based control. Drama 3's promise is that you write something like "she says it quietly, almost embarrassed, then speeds up when she gets to the price" and the model interprets it. That looks more like directing an actor than programming a synthesizer.

The difference is practical:

table · 3 cols
Tags / SSMLPlain-language direction
Learning curveYou must learn the vocabularyWrite the way you would brief a voice actor
PrecisionHigh if you know the tagsDepends on how well the model interprets prose
ReproducibilityDeterministic markupProse can be interpreted variably
Non-technical usersDifficultAccessible
ToolingEasy to validateHarder to lint

The trade-off is control versus accessibility. Natural-language direction is more accessible but may be less repeatable; the true test is whether the same instruction yields the same delivery across several runs.

The three features, and how to test each

1. Mid-sentence voice shifts

Real speech changes within a sentence: a joke that turns sincere, a whisper that rises into a shout. Most TTS handles this by splitting text into separate segments, which breaks flow. If Drama 3 can shift inside a sentence, you avoid stitching artifacts.

Test it: "I told them it was fine, [then, quietly, breaking] and it wasn't." Try three variants and listen for smooth transitions, not a splice.

2. Multi-character scenes in one generation

Dialogue usually needs one call per speaker, then manual timing. A single-pass scene means the model manages turn-taking, overlap and consistency.

Test it: a three-line exchange with two characters interrupting each other, with different ages and accents. Check that voices stay distinct across the whole scene and that interruptions feel natural, not gapped.

3. Single-word fixes

This is the feature editors will love. In current workflows, a mispronounced brand name means re-rolling the entire clip and hoping the rest survives. A one-word patch keeps the parts you approved.

Test it: generate a paragraph with a product name and an unusual surname, then fix only those words. Confirm that loudness, pace and breath match on either side of the patch.

How it compares to other 2026 voice models

The same 48 hours produced two other major voice announcements, and the field is moving fast.

  • Google's Gemini 3.8 Flash TTS adds prompt-based voice design, 2,000+ voices, 100+ languages and 30-second voice replication with consent verification. See Gemini 3.8 Flash TTS.
  • OpenAI's ChatGPT Voice now works with plugins and GPT-6 models for conversation, not TTS production. See ChatGPT Voice plugins.
  • Cartesia Sonic 3.6 was covered against Artificial Analysis benchmarks in our Sonic 3.6 post.
  • Regional and open models: Sarvam Bulbul v4 for Indic languages, Rumik OSS-1 for open-source Indic emotion, and locally runnable options like Kokoro and Miso One.

Where Drama 3 differentiates is not language count or price but direction quality and editability. If you produce dialogue-heavy content, that is exactly the axis that matters.

What people are asking

"How do I get a key?" Reply DRAMA on the announcement thread, then check DMs. Access is manual, which suggests a limited preview.

"What does it cost?" Unpublished. Do not plan a budget around a guess.

"Is it open source?" Nothing suggests so. Fish Audio does maintain open-source speech projects, but Drama 3 is described as an API preview. For self-hosted alternatives see Voicebox and Kokoro.

"Can I clone voices?" Not addressed. Assume nothing about consent or watermark policy.

A practical evaluation checklist

Before you commit, run this on your real scripts:

  1. Emotion sweep. The same sentence in five emotions; grade whether each is recognizable.
  2. Repeatability. Run the same direction five times. Measure variance in pacing and intensity.
  3. Long-form stability. A five-minute passage; listen for drift in voice identity.
  4. Numbers and names. Dates, prices, brand names, acronyms.
  5. Multi-speaker scene. Two or more speakers with overlap.
  6. Single-word patch. Fix, compare loudness and breath continuity.
  7. Latency and cost. Time to first audio and cost per finished minute.
  8. Safety and rights. Ask for commercial-use terms, voice consent policy and watermarking details.

Write down results per criterion so you can compare vendors fairly. Our audio and voice prompt library has prompts you can adapt for the direction tests.

When Drama 3 could matter most

  • Game and animation dialogue, where characters need consistent identity and quick retakes.
  • Audiobooks and narration, where a single-word fix saves hours of re-recording.
  • Ads and explainers, where a client asks for "a bit warmer" and you need to change delivery without redoing the take.
  • Interactive agents that need expressive replies, once latency numbers exist.

Sample direction prompts to try once you have a key

Write these like notes to an actor, then compare the results across runs.

  • "Cheerful customer-support voice, mid-thirties, calm pace. On the refund line, get softer and slower."
  • "Two characters: a tired night-shift nurse and an anxious visitor. The visitor interrupts twice. Keep both voices distinct."
  • "Read the disclaimer quickly and flatly, like small print, then return to a warm tone for the sign-off."
  • "Same sentence three ways: proud, embarrassed, sarcastic."
  • "Fix only the word 'Nakamura' so it is pronounced NAH-kah-moo-rah; keep everything else."

Bottom line

Drama 3 is an interesting product direction: direct voices like an actor, edit takes like text. The three abilities are exactly the ones production teams miss. But this is a gated preview with no public pricing, benchmarks or policy details, so the correct move is to request access, test on your own scripts, and wait for independent measurements before rebuilding a pipeline.

Details reflect Fish Audio's September 23, 2026 announcement and reporting on the following day. The model is a preview and may change.

Related reading

  • Fish Audio $52M seed and S2.1 Pro
  • Gemini 3.8 Flash TTS: voice design in 100+ languages
  • ChatGPT Voice gets plugins, GPT-6 and Work
  • Cartesia Sonic 3.6 on Artificial Analysis
  • Sarvam Bulbul v4: emotional Indic TTS
  • Kokoro: local CPU TTS guide
  • Voicebox: open-source AI voice studio
  • Official: Fish Audio documentation
Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 24, 2026

Gemini 3.8 Flash TTS and Flash-Lite TTS: Voice Design, 2,000+ Voices and 30-Second Voice Cloning

Google DeepMind released two new text-to-speech models in the Gemini API and AI Studio: Flash for creative direction and character design, Flash-Lite for cost-efficient scale. They add prompt-based voice design, 2,000+ ready voices, 100+ languages, voice replication from 30 seconds, and a first-place claim on Hume AI's Voice Design Benchmark.

Sep 8, 2026

Rumik-OSS-1: The First Open-Source Indic TTS With Emotion

Rumik Intelligence released Rumik-OSS-1 on September 8, 2026 — a 3-billion parameter open-weight TTS model built on Cohere's Tiny Aya Fire, claiming to be the first open-source Indic speech model with expressive emotion and parametric voice control across 22 languages.

Sep 7, 2026

VoiceStudio: The Open-Source, Fully-Local ElevenLabs Alternative

VoiceStudio (formerly OmniVoice Studio) is an open-source, AGPL-3.0 desktop app that clones voices, dubs video into other languages, dictates system-wide, and produces audiobooks — all running locally, with 16 TTS engines and 11 ASR engines to choose from.