DeepSeek-V3.2 in full: 61 layers, and what to read next

Putting it together: 61 layers, 671.9 billion weights of which about 37.5 billion are used per token, the cost of the cache, and the papers worth reading next.

This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.

A layer in a box

Zooming out: a layer on the stream

Stacking 61 layers

And there is more

Papers and sources