The KV cache, multi-query and grouped-query attention
Why a language model caches keys and values while it generates, how large that cache gets (gigabytes at 128K tokens), and the first ways to shrink it.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.
Generating one token at a time
- The block runs once per token: the last column = prediction for what comes next (here a full stop)
Appending the new token
- Token appended; run again on a sequence one longer
- Input table has one new column
What changes: the mask
- Because of the mask, old tokens never look ahead → their columns are unchanged (gray)
- Only the new token adds: one new q, k, v; a new map row; a new result column
What is still needed
- To predict the next token: new query vs every old key; result = blend of every old value
- Everything else already did its job: black it out
The KV cache
- Keep keys and values of every token: the KV cache
- New token: compute its q, k, v; read the cache; append its k, v
- Cache: 2 × 128 numbers per token
- FLOPs: all 8 tokens ≈ 7.5 G → new token only ≈ 0.94 G, 8× less (map counted in full)
- The saving grows with every token
Next token: the cache grows
- Next token, same story: only new things computed; cache grows by a column
All heads: a bigger cache
- One head of many: every head has its own K, V
- 128 heads → 32,768 numbers per token, in this layer
Long context: the price
- At 128 k tokens: ≈ 4.3 B numbers ≈ 9 GB in one layer; the model has 61 layers
- Much of what follows shrinks this cache
- FLOPs ≈ 9.5 G per token in this layer, mostly attention over the context
Sharing K and V (MQA)
- Idea 1 (blunt): one K and V shared by all heads (multi-query attention)
- Cache 128× smaller, but quality drops: heads can’t specialise
- Real designs are compromises on this
Groups of heads (GQA)
- Softer: groups of heads share K, V; here 8 groups of 16 (grouped-query attention)
- Cache 8× the shared version, 16× smaller than full; quality stays close
- Many open models use it
Papers and sources
- Vaswani et al. (2017): Attention Is All You Need (queries, keys and values)
- Noam Shazeer (2019): Fast Transformer Decoding: One Write-Head Is All You Need (multi-query attention)
- Ainslie et al. (2023): GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints (grouped-query attention)