Attention step by step: queries, keys, values and heads
Self-attention built up from a single token: queries, keys and values, the attention map and causal mask, softmax, the output matrix and many heads, with DeepSeek-V3.2's dimensions.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão . Open the page with scripts on to play the slides, one key press per idea.
Zooming into the first block Zoom into the first block: what’s inside?
The problem: tokens need context Input table, pick “dog” A word alone is ambiguous: “bat” = animal or something you swing, same numbers in both sentences Meaning must come from the neighbours: enrich each token with associated tokens → the transformer
Query: what am I looking for? Let a token ask a question about what it’s looking for Everything learned, even the questions: a question = a vector, length our choice Here: a question for “dog”
The query as a dense layer Same thing as a neural network, like the first ones Token in (7,168 numbers), question out (128) Each output = a neuron wired to all 7,168 inputs: a dense layer 7,168 × 128 = 917,504 weights (only a few drawn)
The query matrix W_Q Q matrix: question length (128) × token length (7,168), learned from random The model learns what to ask New counter: FLOPs per token for this layer
Making room for the answers (Question calculation moves up to make room for the answers)
Keys: what each token offers How can tokens answer, efficiently and in parallel? Each token makes a vector saying what it’s good at answering: its key
The key matrix W_K The key comes from the K matrix: same shape as Q, learned the same way
Every token at once Q and K are the same for every token: do all tokens at once Whole table × Q → questions; × K → keys; “dog” is one column One matrix multiply: what a GPU does in parallel FLOPs jump: seven tokens’ work at once
Q · Kᵀ: the attention map Questions down the side, keys along the top: each cell = a dot product (big when vectors align) Result: the attention map; row = asker, column = answerer Highlighted row: “dog” (illustrative values)
The causal mask Generating text: no peeking at later tokens → mask Above the diagonal blocked out (−∞) “dog” sees “the”, “old”, itself; not “sleeps”
The size of the map Map is context × context: 7×7 here Double the context → map 4× bigger: cost grows with the square Where much of the rest of the talk is going
Values: what to take The map says where to look, not what to take Third vector: the value, the information a token has to give Back to the input table, still “dog”; the map shrinks into the corner
The value vector Token vector becomes a value vector
The value matrix W_V Third learned matrix, V: value length × token length; random at the start, learned
A table of values Again for every token at once: a table of values
Applying the map (softmax) Softmax on the scores → weights between 0 and 1 summing to 1; masked cells = 0 Map on its side, multiply the values: each result = a blend of the values it attends to “dog”: mostly “old”, a little of itself An enriched token
The result: 128 numbers Result: 128 numbers per token; the stream needs 7,168 So turn it back into a full-length vector
Back to 7,168 numbers 128 numbers → 7,168 again
The output matrix W_O Learned output matrix: 7,168 rows × 128 columns; random at the start
A table of outputs For every token at once: a 7,168-tall table Same shape as the input → can be added back onto the stream
The whole attention operation Whole attention on one screen: input → Q, K, V; Q·K → map; map blends the values; output matrix → input shape All tokens at once
One attention head Box closes round one head: its Q, K, V, the map, its slice of the output matrix (7,168 × 128) Input table left, output right, stream above
Many heads (128) Real models run many heads (128 here), each with its own matrices: different relationships Outputs added together (Σ), sum added to the stream (= concatenate + one big output matrix) DeepSeek-V3.2 dimensions as the baseline; later we change them
Papers and sources Vaswani et al. (2017): Attention Is All You Need (queries, keys and values)