DeepSeek's multi-head latent attention (MLA)
Caching a small latent vector instead of full keys and values, and absorbing the up-projections into the query and output so nothing has to be rebuilt.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão . Open the page with scripts on to play the slides, one key press per idea.
Undoing the sharing Latent attention: first undo the sharing K, V back inside; every head has its own cache again
Idea: cache the input instead Idea: cache the input (7,168 per token), compute K, V when needed vs 32,768 for all heads’ K, V: ≈ 4.6× smaller Cost: recompute the whole context’s K, V every token: ≈ 62 T ops, 6,000+× more
Latent attention (MLA): the KV latent Cache something smaller: squeeze the input into a latent, 512 numbers per token K, V rebuilt by per-head up-projections (128 × 512): 128 heads → 16,384 numbers of keys, same for values Only the latent is cached (real design adds 64 position numbers → 576) Rebuild still ≈ 4.4 T ops per token; the next two steps remove it
MLA: the query latent Same trick on the queries (never cached): W_Q → down to 1,536, up per head Shrinks weights and training memory The indexer (next) will read this query latent
MLA: reordering the matrices (keys) Score = query × key; key = W_UK × latent; multiplication order is free Multiply the query by W_UKᵀ first: a 512-number query in latent space Score the cached latents directly; no keys rebuilt Position part of the keys: small separate path, not drawn ≈ 2.2 T ops (halved)
MLA: reordering the matrices (values) Values the same: weighted sum of latents first, apply W_UV once to that result Used when generating; reading a prompt still rebuilds K, V Paper: could merge into Q and O; released code keeps them separate Rebuild gone: trillions → ≈ 35 B ops per token
Papers and sources DeepSeek-AI (2024): DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model (multi-head latent attention)