DeepSeek's multi-head latent attention (MLA)

Caching a small latent vector instead of full keys and values, and absorbing the up-projections into the query and output so nothing has to be rebuilt.

This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.

Undoing the sharing

Idea: cache the input instead

Latent attention (MLA): the KV latent

MLA: the query latent

MLA: reordering the matrices (keys)

MLA: reordering the matrices (values)

Papers and sources