explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • What Artificial Analysis actually measured
  • The Suleyman claim does not match this board
  • Price versus quality: who actually wins the invoice
  • What this means for builders: STT as the cheap ear
  • What people are asking
  • Honest limitations
  • The takeaway
  • Related on explainx.ai
← Back to blog

explainx / blog

MAI-Transcribe-2-Streaming: #1 Real-Time STT, Not the Cheap Win

Microsoft, Speech-to-Text, Voice AI, Voice Agents, Model Launch

MAI-Transcribe-2-Streaming leads AA-WER at 2.5% WER / 0.13s. Streaming costs $9/hr — more than ElevenLabs, not 60% cheaper.

Oct 2, 2026·14 min read·Yash Thakker
add explainx.ai
go deep
MAI-Transcribe-2-Streaming: #1 Real-Time STT, Not the Cheap Win

On October 1, 2026, Artificial Analysis put Microsoft AI's MAI-Transcribe-2-Streaming at the top of AA-WER Streaming: 2.5% word error at 0.13 seconds after end of speech, first of 38 streaming speech-to-text systems. That is a clean win on the same board where Muse Voice Transcribe sat at 3.1% a month earlier and Gemini 3.5 Transcribe Live sat at 4.0% at launch.

Mustafa Suleyman, CEO of Microsoft AI, called it the most accurate real-time transcription model and told builders to come build agents. The accuracy half of that sentence matches the leaderboard. The marketing half — 55% faster and 60% cheaper than ElevenLabs — does not. This post treats Artificial Analysis as the source of truth, then answers the questions people actually asked on X: why this is not Windows speech-to-text, whether the weights are open, when it shows up in Copilot, and whether streaming STT at $9 an hour is "too cheap to meter" or just another API line item.

TL;DR

table · 2 cols
QuestionDirect answer
What is it?Microsoft AI's streaming speech-to-text model, joining non-streaming MAI-Transcribe-2.
Final WER / latency?2.5% WER at 0.13s after end of speech. #1 of 38 on AA-WER Streaming (Oct 1, 2026).
First partial?2.5% WER at 0.12s. Same error, almost no extra wait for a usable first draft.
Streaming price?$0.54/hour = $9.00 per 1,000 minutes. Level with Gemini 3.5 Transcribe Live.
Non-streaming price?$0.10/hour = $1.67 per 1,000 minutes. The cheap SKU if you can wait.
Windows built-in STT?No announced replacement. Leaderboard win ≠ inbox update.
Open weights?Not announced. Hosted API quality, not a local checkpoint.
In Copilot / Dragon Copilot today?No ship date in this result. July's MAI-Transcribe-1.5 story was a different generation.
Use it as a voice-agent ear?Yes, if you want lowest AA-WER at ~130ms and can pay $9/1k min. Pair with a decision model for intent, not a full speech-to-speech hop.
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What Artificial Analysis actually measured

AA-WER Streaming is not a vendor slide. It is Artificial Analysis's composite word error rate on streaming transcripts, split into Final Transcript (what you keep after the model decides speech has ended) and First Partial (the first committed draft you can act on). Latency is time after end of speech, not time from the first phoneme. Lower WER and lower seconds both win; the interesting models sit in the bottom-left of that scatter.

Final Transcript — October 1, 2026

table · 3 cols
ModelAA-WER (Final)Time after end of speech
MAI-Transcribe-2-Streaming2.5%0.13s
Grok Voice Transcribe 2.02.7%0.49s
Muse Voice Transcribe3.1%0.16s
Cartesia Ink Preview3.1%0.11s
ElevenLabs Scribe v2 Realtime3.6%0.14s

MAI is more accurate than every named neighbor on this board. It is slightly slower than Cartesia Ink Preview (0.13s vs 0.11s) and essentially tied with ElevenLabs on clock time (0.13s vs 0.14s) while beating ElevenLabs by a full percentage point of WER. Grok is close on error and almost 4× slower to finalize. Muse, which owned this chart in early September, is now a third of a point worse and 30 milliseconds slower.

First Partial — same day, same board

table · 3 cols
ModelAA-WER (First Partial)Time after end of speech
MAI-Transcribe-2-Streaming2.5%0.12s
Muse Voice Transcribe3.6%0.13s
ElevenLabs Scribe v2 Realtime3.6%0.13s
Grok Voice Transcribe 2.03.4%0.49s
Cartesia Ink-24.0%0.07s

First Partial is the number that matters for a voice agent that starts planning a reply before the user has fully stopped. MAI keeps 2.5% WER at 0.12s. Cartesia Ink-2 still wins the sprint (0.07s) and pays for it with 4.0% error — faster, worse. Muse and ElevenLabs land at 3.6% / 0.13s. Grok's first partial is still waiting half a second.

If you only remember one comparison: MAI is the accuracy leader at interactive latency, not the latency leader at any price. Cartesia still owns "first token of text as fast as possible." MAI owns "the words you show the user are the ones you meant."

The Suleyman claim does not match this board

Suleyman said the model is 55% faster and 60% cheaper than ElevenLabs, and that it is the most accurate real-time transcription model, and that people should come build agents on it.

Three of those clauses can be true at once only if "ElevenLabs" means something other than Scribe v2 Realtime on Artificial Analysis.

On the board we have:

  • Price. MAI streaming is $9.00 per 1,000 minutes. ElevenLabs Scribe v2 Realtime is $6.50. MAI is more expensive, not 60% cheaper. A 60% cut from $6.50 would be $2.60. MAI is $9. That is roughly 38% more, not 60% less.
  • Speed. Final transcript is 0.13s vs 0.14s. That is about 7% faster, not 55%. First Partial is 0.12s vs 0.13s — same story.
  • Accuracy. 2.5% vs 3.6% Final WER is a real gap. "Most accurate real-time transcription" is the claim the leaderboard supports.

This is marketing versus leaderboard, not a rounding error. Possible reconciliations we cannot verify from public AA numbers: a different ElevenLabs SKU (batch, studio, or an older realtime tier), a Microsoft-internal latency probe that is not AA's end-of-speech clock, or a bundle price that mixes non-streaming MAI ($1.67 per 1,000 minutes) with a streaming competitor. Non-streaming MAI at $1.67 versus streaming ElevenLabs at $6.50 is about 74% cheaper — still not 60%, and it compares batch to realtime, which is not how you price a voice agent.

Until Microsoft publishes the baseline SKU, treat 2.5% / 0.13s / $9 as the number you put in a design doc, and treat 55% / 60% as a campaign line. explainx.ai's read is that the accuracy win is real and the price/speed slogan is not reconcilable with the same dataset that made the model famous.

Price versus quality: who actually wins the invoice

Streaming speech-to-text prices on the same October 1 snapshot:

table · 3 cols
ModelStreaming pricevs MAI
MAI-Transcribe-2-Streaming$9.00 / 1,000 min ($0.54/hr)—
Gemini 3.5 Transcribe Live$9.00Level
ElevenLabs Scribe v2 Realtime$6.50MAI is +38%
Deepgram Flux$6.50MAI is +38%
Cartesia Ink-2$4.00MAI is more than 2×
Muse Voice Transcribe$3.00MAI is 3×

Non-streaming MAI-Transcribe-2 is a different product: $0.10/hour, $1.67 per 1,000 minutes. That is where Microsoft is cheap. If you are captioning podcasts, doing offline meeting notes, or batching call-center archives, do not buy the streaming SKU. If you are running a live agent, you are on the $9 line.

Gemini 3.5 Transcribe already taught this lesson in August: Google's live hop is the expensive one, the pre-recorded hop is the cheap one, and mixing the two in a tweet is how you get a fake bargain. MAI is repeating the pattern with a better WER.

What you buy at $9 is lowest measured streaming error at human-conversation latency. What you do not buy is the cheapest ear. Muse at $3 and 3.1% / 0.16s is the "good enough and 3× cheaper" default. Cartesia at $4 is the "faster first partial, worse WER" default. ElevenLabs and Deepgram sit in the middle on price and lose on AA-WER.

For a voice agent that talks for an hour a day per seat, $9 per 1,000 minutes is $0.54 per hour of audio. That is not "too cheap to meter" unless your agent is a toy. A contact-center seat at four hours of talk time a day is about $2.16/day in STT alone, before the LLM, TTS, telephony, and tool calls. Meter it. Put a kill switch on silence and on hold music. Do not stream the parking lot.

What this means for builders: STT as the cheap ear

A voice agent is not one model. It is an ear, a brain, and a mouth. Hugging Face speech-to-speech still writes that pipeline as VAD → STT → LLM → TTS. GPT-Live-1 collapses ear and mouth into one full-duplex speech-to-speech call and leaves reasoning to a backend model. MAI-Transcribe-2-Streaming is a bid to own only the ear, at frontier WER, and let you pick the rest.

That is the useful architecture for most production agents in 2026:

  1. Streaming STT commits partials so the agent can start planning at ~120ms after the user stops, not after a 400–500ms Grok-style finalize.
  2. A decision model (or a tiny classifier) answers the questions that should never wait on a chat completion: is this barge-in, is this end-of-turn, is this the billing intent or the cancel intent, should we keep listening.
  3. An LLM writes the reply, calls tools, and handles the long tail.
  4. TTS speaks. Or you skip TTS and type, if this is a dictation surface.

The ear should be cheap relative to the brain. MAI's streaming SKU is not cheap relative to Muse or Cartesia. It is cheap relative to paying GPT-Live-1 or another speech-to-speech model to listen when you only needed text. If your agent already has a tool-calling harness, buying a specialist STT and a specialist decision head is usually less invoice and less lock-in than buying a duplex voice model that also wants to be the personality.

Practical split:

table · 3 cols
WorkloadBuy this earWhy
Live agent, English-heavy, errors are expensive (payments, medical intake, legal)MAI-Transcribe-2-Streaming2.5% WER at 0.13s is the board's best finalize.
Live agent, cost-sensitive, 3.1% WER is fineMuse Voice TranscribeSame latency band, one-third the streaming price, plus native diarization if you need speakers.
Live agent, interrupt-heavy UI, first glyph winsCartesia Ink-2 / Ink Preview0.07–0.11s; accept higher WER and correct downstream.
Offline notes, podcasts, archivesMAI-Transcribe-2 non-streaming$1.67 / 1,000 min. Do not stream a file you already have.
Open pipeline, on-prem, or swap-STT experimentsHF speech-to-speech + local STTWeights you can move. Not 2.5% AA-WER.
Code-mixed Indian-language trafficBenchmark Sarvam Voice AgentsAA-WER is an English-weighted composite. Do not import a 2.5% as a Hindi-English number.

A copy-paste eval loop you can run this week, regardless of vendor:

text
For each 60s clip in your production sample:
  1. Stream audio; record t_end_of_speech, t_first_partial, t_final.
  2. Store first_partial_text and final_text.
  3. Score WER against a human transcript.
  4. Score "intent flip": did first_partial route to the same tool as final?
  5. Reject a vendor that wins WER but flips intent on first_partial.

The last line is the one leaderboards skip. A 2.5% WER model that still says "cancel" instead of "transfer" on the first partial will fire the wrong tool. Pair STT with a decision model that refuses to act until confidence on the intent schema is high, even if the transcript already looks clean.

What people are asking

The X thread under Suleyman's post was not about WER decimals. It was about Microsoft's own surfaces.

Why isn't this Windows speech-to-text?

Because a Foundry or Microsoft AI API ranking #1 is not a Windows inbox. Built-in Speech Recognition, Live Captions, and Voice Access are OS features with offline constraints, accessibility contracts, and a decade of enterprise GPOs. Shipping MAI-Transcribe-2-Streaming as "the Windows ear" would mean on-device or at least edge-local weights, a privacy story, and a servicing plan. None of that is in the October 1 leaderboard card.

Until Microsoft says otherwise, treat this as cloud STT you call, the same way you call Gemini Live or Muse. If you were hoping to rip out a third-party dictation app on a locked-down Windows laptop, this announcement does not do that.

Are the weights open?

No announcement as of October 2, 2026. MAI specialists have mostly shipped as hosted Foundry models — the same pattern as the July hill-climbing writeup for Copilot and Excel. Open weights would change the story (local agents, air-gapped call centers, Hugging Face backends). Closed weights keep this in the same bucket as Gemini 3.5 Transcribe: great API, no laptop story.

If open is the requirement, the current builder path is still Hugging Face speech-to-speech with a swappable local STT, not MAI.

When does this hit Copilot?

Unknown. July's Dragon Copilot note was MAI-Transcribe-1.5 and a Microsoft-reported 50% relative cut in transcription and language-ID error across most of 58 languages — a product metric, not AA-WER. This October result is MAI-Transcribe-2-Streaming on an independent English-weighted streaming board. Do not collapse the two into "Copilot just got 2.5% WER."

Suleyman's "come build agents" is the tell. The first customer is developers on an API, not the Word ribbon. Copilot will get it when the hill-climbing machine decides the product evals moved. That is consistent with how Microsoft already routes MAI versus partner models: ship the specialist when the harness says it won, not when the public leaderboard says it won.

Is it "too cheap to meter"?

No. $9 per 1,000 streaming minutes is mid-to-high on this board. The phrase fits non-streaming at $1.67, or Muse streaming at $3, not the model that just took #1. If your CFO heard "60% cheaper than ElevenLabs" and approved a fleet-wide cutover, send them the $9 vs $6.50 row before you cut over.

Honest limitations

  • AA-WER is a composite, not your audio. The Gemini launch post already warned that Google's 2.6% non-streaming on AA's mix ballooned to 5.04% on public FLEURS. Treat 2.5% as a ranking, then measure your domain, accent, and channel.
  • No diarization story in this result. Muse's September pitch was ASR + speakers + endpointing in one model. MAI's October pitch is WER and latency. If you need "who said it," ask for a separate metric before you rip out a diarization hop.
  • No language mix in the tweet. July's 1.5 claim was multilingual. This 2.x streaming #1 is an AA-WER Streaming rank. Do not assume 58-language parity.
  • No open weights, no Windows SKU, no Copilot GA date. The gaps people asked about are still gaps.
  • Suleyman's 55%/60% line is unverified. We will update if Microsoft publishes the ElevenLabs SKU and clock it used.
  • Vendor access and rate limits are not in the AA card. Being #1 on a benchmark is not the same as being generally available in your tenant.

The takeaway

MAI-Transcribe-2-Streaming is the new accuracy leader on Artificial Analysis's streaming speech-to-text board: 2.5% WER, 0.13s to final, 0.12s to first partial, $9 per 1,000 minutes. That is the sentence to keep. Use it as the ear of a voice agent when errors are expensive and you already have a brain — a decision model or an LLM — that should not also be listening to raw audio.

Do not use the tweet as a price card. Do not wait for Windows to swallow it. Do not confuse it with last summer's Dragon Copilot 1.5 metric. And do not pay streaming rates for files you could batch at $1.67.

Related on explainx.ai

  • Gemini 3.5 Transcribe: what builders actually get — the $9/1k-min live SKU this one ties on price
  • Muse Voice Transcribe — previous AA-WER Streaming leader, still the cheap 3.1% option
  • GPT-Live-1 in the API — speech-to-speech instead of STT-as-ear
  • Hugging Face speech-to-speech voice agents — open VAD → STT → LLM → TTS when weights matter
  • Microsoft MAI hill-climbing — MAI-Transcribe-1.5 in Dragon Copilot, and why product evals lag public boards
  • What are decision models? — the intent layer to sit behind a streaming ear
  • Sarvam Voice Agents — do not import an English AA-WER into code-mixed traffic

Sources: Artificial Analysis AA-WER Streaming snapshot, October 1, 2026 (Final Transcript and First Partial; streaming and non-streaming prices as listed above) · Mustafa Suleyman public remarks on MAI-Transcribe-2-Streaming, October 2026


WER, latency, and prices in this post reflect Artificial Analysis as of October 1, 2026 and Microsoft AI public comments as of October 2, 2026. Leaderboards, SKUs, and product surfaces move — re-check AA and your Microsoft tenant before you lock a vendor.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Sep 2, 2026

Muse Voice Transcribe: Meta's Real-Time ASR, Diarization, Endpointing in One Model

Meta Superintelligence Labs shipped Muse Voice Transcribe on September 1, 2026 — a single model doing streaming speech-to-text, speaker diarization, and endpointing natively, with benchmark numbers that beat the ElevenLabs, Deepgram, Cartesia, and AssemblyAI pipelines builders currently stitch together by hand.

Aug 27, 2026

Gemini 3.5 Transcribe: What Builders Actually Get

Google launched Gemini 3.5 Transcribe on August 26, 2026 with two model IDs, 85+ language auto-detection, and function calling inside the transcription call. explainx.ai breaks down the real pricing, the smart-vs-verbatim trade-off that decides whether it is legal for your use case, and where it loses to the incumbents.

Sep 11, 2026

GPT-Live-1 Is Now in the API: OpenAI Opens Its Voice Model to Builders

OpenAI Developers announced GPT-Live-1 is now available in the API on September 10, 2026 — the full-duplex speech-to-speech model that has powered ChatGPT Voice since July finally ships as infrastructure builders can call directly, paired with whatever reasoning model and tool-calling harness they choose.