Tokens and embeddings: how text becomes numbers
How a language model turns text into tokens and each token into a learned vector: the tokenizer, DeepSeek's vocabulary of 129,280 tokens, and the embedding table.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão . Open the page with scripts on to play the slides, one key press per idea.
A vote over tokens Same vote, but the output neurons are tokens (think words); real DeepSeek-V4 ids scroll past; ␣ = leading space Vocab sizes: GPT-2 ≈ 50 k, Llama 3.1 128 k, DeepSeek-V4 129,280, gpt-oss 201,088, Gemma 2 256 k Bigger vocab = fewer tokens per sentence, bigger first and last layers (no honest number for Claude 5.5 or GPT-6: tokenizers unpublished) Input side: text → tokens → numbers; type in the box (digits in groups of ≤ 3; Chinese in whole pieces) The question: given what came before, which word next? (mouse eats the → cheese)
Tokens in, tokens out Both ends changed shape: neurons became arrays In: tokens as numbers; out: one column, a probability per vocabulary entry (illustrative, 129,280) A layer = matrix × vector: what a GPU does Middle unknown; to see it fully I trained a tiny model on a 60-token toy language
The input: an array of tokens Open the box at the input: the real tokens of what you typed The array tips over into the distance, like the digit image
Why not feed in token numbers? Why not feed in token numbers? They mean nothing: 3,072 isn’t “a bit more” than 3,071 We want the network to learn each word’s meaning So make room under each token: an empty column for its representation
The embedding table Embedding table: one column per token (129,280), each 7,168 numbers (DeepSeek’s hidden size) Random at the start, learned like any weight 129,280 × 7,168 ≈ 927 M weights, 1.9 GB, before doing anything Illustrative numbers; thousands of columns skipped
Looking up a column Token number = column number: copy that column under the token Same token twice → same column: a lookup, not a calculation Result: a grid, one column per token, a learned image of the text
Papers and sources Sennrich, Haddow & Birch (2015): Neural Machine Translation of Rare Words with Subword Units (byte-pair encoding) Bengio, Ducharme, Vincent & Jauvin (2003): A Neural Probabilistic Language Model (learned word vectors: embeddings)