Tokens and embeddings: how text becomes numbers

How a language model turns text into tokens and each token into a learned vector: the tokenizer, DeepSeek's vocabulary of 129,280 tokens, and the embedding table.

This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.

A vote over tokens

Tokens in, tokens out

The input: an array of tokens

Why not feed in token numbers?

The embedding table

Looking up a column

Papers and sources