A Deep Seek into LLM Architecture

An animated, step-by-step talk that builds up to the architecture of DeepSeek's language models from a single neuron: layers, convolutions, ResNet, tokens and embeddings, attention, the KV cache, multi-head latent attention, sparse attention and mixture of experts. Every number is drawn from the published papers.

By Ruben Galvão. The talk is played one key press at a time (arrow keys, space, or swipe), and each chapter below has its own page with the text of its steps and the papers it draws on.