Attention step by step: queries, keys, values and heads

Self-attention built up from a single token: queries, keys and values, the attention map and causal mask, softmax, the output matrix and many heads, with DeepSeek-V3.2's dimensions.

This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.

Zooming into the first block

The problem: tokens need context

Query: what am I looking for?

The query as a dense layer

The query matrix W_Q

Making room for the answers

Keys: what each token offers

The key matrix W_K

Every token at once

Q · Kᵀ: the attention map

The causal mask

The size of the map

Values: what to take

The value vector

The value matrix W_V

A table of values

Applying the map (softmax)

The result: 128 numbers

Back to 7,168 numbers

The output matrix W_O

A table of outputs

The whole attention operation

One attention head

Many heads (128)

Papers and sources