Grouped Query Attention (GQA) is a variant of multi-head attention where multiple query heads share a single key-value head, reducing KV cache size and memory bandwidth during inference while retaining most of multi-head attention's quality. It offers an effective middle ground between the full KV overhead of multi-head attention and the aggressive sharing of multi-query attention.