The residual stream: the backbone of a transformer
The grid of token vectors that runs through a transformer, which every block reads from and adds to, ending in the next-token prediction.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.
The input as a grid
- Hide the table; slide the grid left
- This grid travels through the whole network: the residual stream
- Starts as embeddings; everything else only adds
The final layer
- Show the ending first: final layer → a score per vocabulary token (illustrative)
- V3.2: norm, one 129,280 × 7,168 matrix on the last column only, softmax (in the sampling code)
- We know what goes in and what comes out; the middle is unopened
The residual stream
- Stream runs from the input table to a same-shape output table
- Between: blocks (question marks) hanging off the stream
- Each reads the stream and adds its result; never replaced
- Opening those blocks is the rest of the talk
Only the last column is used
- The final layer reads only the last column
- Earlier columns already fed it through the blocks: that’s how context arrives
- New token joins the row, and it all runs again