DeepSeek Sparse Attention and the lightning indexer

A cheap lightning indexer scores every cached token so that attention only reads the best 2,048, cutting the memory read for long contexts.

This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.

Reading the whole cache is costly

The lightning indexer

The indexer as a box

Top-k: sparse attention

Papers and sources