DeepSeek Sparse Attention and the lightning indexer
A cheap lightning indexer scores every cached token so that attention only reads the best 2,048, cutting the memory read for long contexts.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.
Reading the whole cache is costly
- Cache is small, but each token still reads all of it: 128 k keys per step
- Most barely matter: a row of scores, a few big
- Guess the scores cheaply → fetch only the few that matter: far less memory bandwidth
The lightning indexer
- Cheap scores: DeepSeek’s lightning indexer: just a smaller attention block
- Own queries per head (from the query latent), own keys shared by its heads in a small cache; heads summed with learned weights (64 × 7,168)
- Fewer heads (64 vs 128), low precision (FP8), vectors 128 long vs 576 in the core
- Cheap enough to run over every cached token
- Counters include its three matrices and key cache
The indexer as a box
- Fold into one box: the indexer returns a score per cached token
- Block moves back up
Top-k: sparse attention
- Keep the best scores: 2,048 (V3.2 paper; V4 keeps 1,024)
- Attention only on those: sparse attention, reads 2 k not 128 k
- V4: indexer over compressed entries, plus the latest 128 tokens always
- Attention ≈ 34 B → ≈ 0.5 B ops; indexer ≈ 2 B; layer ≈ 37 B → ≈ 3 B
Papers and sources
- DeepSeek-AI (2025): DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models (DeepSeek Sparse Attention, lightning indexer)