Diarization is separate from recognition: a model can transcribe every word correctly and still attribute them to the wrong person, which is why vendors publish speaker limits alongside word error rate. Practical ceilings stay low — Google's Gemini 3.5 Transcribe supports attribution for up to three speakers with word-level timestamps and labels anything beyond that experimental — so multi-party meeting products usually still need per-channel audio or a downstream correction pass.