DeepSeek-V3.2 in full: 61 layers, and what to read next
Putting it together: 61 layers, 671.9 billion weights of which about 37.5 billion are used per token, the cost of the cache, and the papers worth reading next.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão . Open the page with scripts on to play the slides, one key press per idea.
A layer in a box Attention block + mixture of experts, each adding to the stream = one layer Draw a box round the two
Zooming out: a layer on the stream Zoom out to the start: stream input → output with blocks The question mark is now a layer: attention, then feed-forward or experts (The one we opened is mixture-of-experts)
Stacking 61 layers V3.2: 61 layers: 3 dense, 58 experts 671.9 B weights from the config (report: 671 B), incl. the two 129,280 tables; not the ≈ 14 B multi-token-prediction layer A token uses ≈ 37.5 B (report: 37 B) Cache at 128 K: 61 × ≈ 185 MB = 11.3 GB; every head with its own K, V: ≈ 524 GB FLOPs per token at 128 K: ≈ 237 B; only ≈ 75 B from the weights (2 × 37.5 B), rest attention + indexer Counters leave out a few small parts (position dims)
And there is more More we don’t have time for: each is a paper Compressed attention (compressed sparse + heavily compressed hybrid; V4) mHC: widens and stabilises the residual stream Engram: memory as a lookup table, a second sparsity next to experts Seek, and you shall find
Papers and sources DeepSeek-AI (2024): DeepSeek-V3 Technical Report