A Deep Seek into LLM Architecture
An animated, step-by-step talk that builds up to the architecture of DeepSeek's language models from a single neuron: layers, convolutions, ResNet, tokens and embeddings, attention, the KV cache, multi-head latent attention, sparse attention and mixture of experts. Every number is drawn from the published papers.
By Ruben Galvão . The talk is played one key press at a time (arrow keys, space, or swipe), and each chapter below has its own page with the text of its steps and the papers it draws on.
The perceptron: one neuron that learns : Start with a single neuron that learns to convert ancient Egyptian cubits to metres: a weight, a loss, backpropagation and gradient descent, which you can try by dragging the weight.Why a chain of neurons adds nothing, and a bend does : Real temperature data from Norwich shows why stacking straight-line neurons never helps, and how a ReLU bend in every neuron lets a network fit a curve.Reading digits: pixels, convolutions and softmax : How a small network reads a handwritten digit: a weight for every pixel, local patches, shared filters (convolution), pooling and a softmax vote, with a live drawing pad.AlexNet: the same recipe at ImageNet scale : AlexNet (2012) is the digit reader scaled up to a thousand kinds of photograph; here a pre-trained image network runs live in the browser on test pictures or your camera.ResNet: why deeper networks got worse, and the shortcut that fixed it : The degradation problem and vanishing gradients, measured on a small network, and how shortcut connections in ResNet give the residual stream that every transformer uses.Tokens and embeddings: how text becomes numbers : How a language model turns text into tokens and each token into a learned vector: the tokenizer, DeepSeek's vocabulary of 129,280 tokens, and the embedding table.The residual stream: the backbone of a transformer : The grid of token vectors that runs through a transformer, which every block reads from and adds to, ending in the next-token prediction.Attention step by step: queries, keys, values and heads : Self-attention built up from a single token: queries, keys and values, the attention map and causal mask, softmax, the output matrix and many heads, with DeepSeek-V3.2's dimensions.The KV cache, multi-query and grouped-query attention : Why a language model caches keys and values while it generates, how large that cache gets (gigabytes at 128K tokens), and the first ways to shrink it.DeepSeek's multi-head latent attention (MLA) : Caching a small latent vector instead of full keys and values, and absorbing the up-projections into the query and output so nothing has to be rebuilt.DeepSeek Sparse Attention and the lightning indexer : A cheap lightning indexer scores every cached token so that attention only reads the best 2,048, cutting the memory read for long contexts.Mixture of experts: the feed-forward block, experts and the router : The feed-forward block, then DeepSeekMoE: 256 small experts plus a shared one, and a router that sends each token to its best eight.DeepSeek-V3.2 in full: 61 layers, and what to read next : Putting it together: 61 layers, 671.9 billion weights of which about 37.5 billion are used per token, the cost of the cache, and the papers worth reading next.