The index reports the percentage of words a streaming model gets wrong against a human reference, plus how many seconds after detected end of speech the first partial and the final transcript arrive. Lower WER and lower latency both win. On October 1, 2026, Microsoft AI's MAI-Transcribe-2-Streaming led the 38-model board at 2.5% final WER in 0.13 seconds. Scores are only comparable inside Artificial Analysis's own audio mix — they are not interchangeable with public sets such as FLEURS.