WER is computed as (substitutions + deletions + insertions) divided by the number of words in the reference transcript, so lower is better and scores above 100% are possible when a model inserts extra words. Because the score depends entirely on which audio it was measured on, WER figures are only comparable across vendors when they share a dataset — Google reported Gemini 3.5 Transcribe at 2.6% non-streaming on Artificial Analysis's dataset mix but 5.04% on the public multilingual FLEURS benchmark, roughly double, from the same model.