Zoom out: all this is the attention block, half a layer; the other half is the feed-forward block
Per token, no mixing: two wide matrices (7,168 → 18,432), one through SiLU (a gate), multiplied cell by cell, a third back to 7,168; adds to the stream
Attention mixes tokens; feed-forward thinks about each one
V3.2: first 3 layers dense, the rest experts
The feed-forward works per token
Darken all but the new token: feed-forward never looks at other tokens
Each column independent through the same three matrices
Position-wise feed-forward (original paper)
Nothing to cache; a new token costs one column
One expert: a narrower block
Mixture of experts: box the three matrices = an expert; make it narrower