The KV cache, multi-query and grouped-query attention

Why a language model caches keys and values while it generates, how large that cache gets (gigabytes at 128K tokens), and the first ways to shrink it.

This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.

Generating one token at a time

Appending the new token

What changes: the mask

What is still needed

The KV cache

Next token: the cache grows

All heads: a bigger cache

Long context: the price

Sharing K and V (MQA)

Groups of heads (GQA)

Papers and sources