explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

On this page

  • TL;DR
  • Hear the two clips Cartesia posted
  • What the two leaderboards actually measure
  • How to try it without burning production
  • Price is not the story. Switching cost is.
  • What to watch next
  • Related on explainx.ai
← Back to blog

explainx / blog

Cartesia Sonic-3.6: #1 on Both Artificial Analysis TTS Boards

Cartesia shipped Sonic-3.6 in beta on August 17, 2026 — #1 on both Artificial Analysis streaming TTS boards, 44 languages, $49 per million characters.

Aug 18, 2026·8 min read·Yash Thakker
Voice AITTSCartesiaArtificial AnalysisSpeech
go deep
Cartesia Sonic-3.6: #1 on Both Artificial Analysis TTS Boards

Cartesia just took both Artificial Analysis streaming TTS boards, three months after it last reset the naturalness race. On August 17, 2026, @cartesia announced Sonic-3.6 as "our most lifelike TTS yet" — 44 languages, beta today, and #1 on both the provider-voice and controlled-voice leaderboards.

That is a real, checkable claim. It is also a preference-arena claim, not a production-readiness claim. This post separates the two, embeds the English and Hinglish clips Cartesia posted in the same thread, and puts the Elo numbers next to price and the rest of 2026's voice stack.

XSource postOpen on X ↗
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR

table · 2 cols
QuestionAnswer
What shipped?Sonic-3.6, Cartesia's latest TTS, in beta
When?August 17, 2026 — three months after Sonic-3.5 (May 4, 2026 snapshot)
API id?Docs currently map it to sonic-preview — beta, can change
Languages?44 (Sonic-3.5 is 42)
Provider-voice Elo?1,282 — #1 (Artificial Analysis)
Controlled-voice Elo?1,123 — #1, ahead of Sonic-3.5 at 1,102
List price?$49 / 1M characters, same as Sonic-3.5 on the AA table
Should prod switch?Not yet. A/B sonic-preview against pinned sonic-3.5

Hear the two clips Cartesia posted

The launch thread is not just charts. Two follow-up posts carry the actual audio: one English, one Hinglish. Both are mirrored locally here so they play without an X widget.

English — natural pauses and filler words (about 3 seconds), from the first video post:

Sound on. Cartesia's own caption is the spec: filler words and pauses, not a studio-clean read.

Hinglish — Hindi/English code-switching (about 10 seconds), from the second video post:

That clip is the more interesting product signal. Seamless code-switching is the thing Indian-language voice stacks — including Sarvam Bulbul V4 — are competing on. A US vendor putting Hinglish in the launch thread is not a coincidence.

What the two leaderboards actually measure

Cartesia said it is #1 "across both the provider and controlled voice streaming leaderboards — best-in-class voices + the best underlying model." Those are two different tests on Artificial Analysis. Mixing them is how people talk themselves into a fake sweep.

Provider voices: each lab's own speakers

This board uses each provider's native voices. Listeners pick the more natural sample in a blind pair. As of this writing:

Artificial Analysis provider-voice TTS leaderboard with Cartesia Sonic 3.6 first at 1282 Elo

table · 4 cols
RankModelEloList $ / 1M chars
1Sonic 3.6 (Cartesia)1,282$49
2Simba 3.2 (SpeechifyAI)1,239$10
3Qwen-Audio-3.0-TTS-Plus (Alibaba)1,237$27.60
4Luna TTS (VUI Labs)1,219$80
5Gemini 3.1 Flash TTS (Google)1,212$18.30
7Sonic 3.5 (Cartesia)1,202$49
11Eleven v3 (ElevenLabs)1,177$100
18Fish Audio S2.1 Pro1,142$15
49Bland Speech v31,065$40

The 80-point gap from Sonic-3.6 to Sonic-3.5 on this board is the "step change" Cartesia is selling. The 43-point gap to Simba 3.2 is smaller than the marketing screenshot suggests once you notice Simba is one-fifth the list price.

Controlled voices: same eight clones for everyone

This board holds the speaker constant — eight cloned voices, four US and four UK — so you are scoring the model, not the vendor's default actor.

Artificial Analysis controlled-voice TTS leaderboard with Sonic 3.6 first at 1123 Elo and Sonic 3.5 second

table · 4 cols
RankModelEloList $ / 1M chars
1Sonic 3.6 (Cartesia)1,123$49
2Sonic 3.5 (Cartesia)1,102$49
3Eleven v3 (ElevenLabs)1,059$100
4StepAudio 2.5 TTS (StepFun)1,049$85
5Realtime TTS 1.5 Max (Inworld)1,047$26
7SpaceXAI TTS1,040$15
14Fish Audio S2 Pro1,004$15
18Fish Audio S2.1 Pro996$15
29Bland Speech v3917$40

Two things jump out:

  1. Cartesia occupies ranks 1 and 2. On the apples-to-apples clone test, 3.6 beats 3.5 by 21 Elo — a real lift, not a new default-voice trick.
  2. DesignArena #1 is not Artificial Analysis #1. Bland Speech v3 launched two weeks earlier as the top model on DesignArena's Audio Realism Bench. On AA's controlled board it sits at 917. Different arena, different listeners, different scripts. Cite the board you actually used.

Elo here is crowdsourced pairwise preference. Artificial Analysis is explicit: users hear the same text from two models and pick which sounds more natural. That is useful. It is not a mean-opinion-score lab, a telephony MOS, or a measurement of first-byte latency. Cartesia still markets Sonic-3.5 as sub-90ms to first byte; the 3.6 post did not publish a new latency number.

How to try it without burning production

Cartesia's model docs still describe Sonic-3.5 as the stable flagship. The API-changes page is where 3.6 actually appears: sonic-preview → "Sonic 3.6 preview," 44 languages.

That mapping is the whole production decision:

table · 3 cols
model_idWhat it doesUse it for
sonic-3.5-2026-05-04Frozen snapshotEvals you can rerun next month
sonic-3.5Tracks latest stable 3.5Production that wants patches
sonic-previewBeta pointer for 3.6Listening tests, not the on-call pager

Clone flow is unchanged from the June 2026 embedding sunset: 3–10 seconds of source audio via POST /voices/clone, then pass a voice ID into TTS. Do not send raw embeddings; those paths already error.

A sane eval for a voice agent this week:

  1. Take 20 real transcripts (including confirmation codes, emails, and mixed Hindi-English if you serve India).
  2. Generate them on sonic-3.5-2026-05-04 and sonic-preview with the same voice ID.
  3. Blind-rank internally. Do not let the person who picked Cartesia last quarter score the files.
  4. Measure time-to-first-audio on your WebSocket path. Naturalness that adds 80ms of jitter loses barge-in.
  5. Keep 3.5 pinned until Cartesia publishes sonic-3.6-YYYY-MM-DD.

Price is not the story. Switching cost is.

$49 / 1M characters is identical to Sonic-3.5 on the Artificial Analysis table. Cartesia did not buy the Elo lead with a discount.

Against the rest of the field explainx.ai already covers:

  • Fish Audio S2.1 Pro is ~$15 / 1M chars and open-weight-adjacent. On the provider board it is 140 Elo behind 3.6. On the controlled board S2.1 Pro is 127 Elo behind. That is a quality gap, not a rounding error — and still not an automatic "pay 3×" decision if your traffic is narration, not a live agent.
  • Eleven v3 is $100 / 1M chars and third on the controlled board. If you stayed on Eleven for cloning tools rather than naturalness, 3.6 is the first Cartesia number that makes a bake-off mandatory.
  • Grok Voice Think Fast is a speech-to-speech path, not a drop-in TTS clone. Do not put it on this Elo table.

The honest worksheet is the same one from the Fish Audio post: your monthly characters, your concurrency, your incident cost if a clone is misused, and the engineering days to swap providers. Sonic-3.6 changes the quality cell. It does not change the consent cell. Fast cloning from a few seconds of audio is still the dual-use problem New York's synthetic-performer disclosure law is trying to get in front of.

What to watch next

  • A dated sonic-3.6-… snapshot. Until that exists, 3.6 is a moving beta.
  • Latency, published. If 3.6 keeps 3.5's sub-90ms first byte, the Elo lead is the whole pitch. If it does not, live-agent teams will stay on 3.5.
  • The two extra languages. Cartesia has not listed which two were added on top of 3.5's 42 in the public docs yet. Ask before you assume they are the ones you need.
  • Whether Simba 3.2 or Qwen-Audio close the provider-voice gap. Those two are already inside ~45 Elo at a fraction of the price.

Related on explainx.ai

  • Fish Audio $52M seed and S2.1 Pro — the cost-and-open-weights counter-pitch
  • Bland Speech v3 — a different #1, on a different arena
  • Sarvam Bulbul V4 — Indian-language expressiveness and why the Hinglish clip matters
  • Grok Voice Think Fast 2.0
  • Voicebox: open-source voice studio
  • New York AI video disclosure law for synthetic performers
  • YC Fall 2026 Requests for Startups — deepfake trust infrastructure

Sources

  • Cartesia — Sonic-3.6 announcement
  • English demo · Hinglish demo
  • Cartesia TTS models (latest)
  • Cartesia API changes / sonic-preview
  • Artificial Analysis — provider-voice TTS
  • Artificial Analysis — controlled-voice TTS

Elo scores, sample counts, and list prices are from Artificial Analysis as of August 18, 2026 and will move as more votes land. Sonic-3.6 is a vendor beta; confirm the current model_id and language list in Cartesia's docs before you ship it.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

Related posts

Jul 29, 2026

Fish Audio Raises $52M and Launches S2.1 Pro Voice AI

One year from zero to $21M ARR and 8M users, Fish Audio closed a $52M seed and publicly launched S2.1 Pro — expressive TTS aimed at ElevenLabs and Cartesia, with a free developer API window and a 50% cost-cut enterprise guarantee.

Aug 10, 2026

New Orleans AI 911 Dispatch: What the First US City Rollout Actually Does

New Orleans has been running Carbyne's AI on its 911 lines since 2023, and in August 2026 confirmed it now triages live emergency calls — filtering duplicates and non-emergencies so human dispatchers focus on the calls that need a human voice. Here's what's actually automated, what stays human, and why 911 dispatch is the highest-stakes human-in-the-loop test case explainx.ai has covered.

Aug 5, 2026

Bland Speech v3: Inside the "Human Speech Engine" Launch

On August 4, 2026, phone-agent company Bland launched Speech v3, a standalone voice model it calls the "world's first Human Speech Engine." The centerpiece is a case study restoring a stroke survivor's voice — here's what the benchmark claim actually rests on and what the launch means for Bland's business.