Multi-Query Attention (MQA) is the extreme case of grouped query attention where all query heads share a single key head and a single value head. This yields maximum KV cache savings and faster inference but slightly lower quality than GQA. It was introduced by Noam Shazeer and is used in models like PaLM and Falcon.