Mixture of experts: the feed-forward block, experts and the router

The feed-forward block, then DeepSeekMoE: 256 small experts plus a shared one, and a router that sends each token to its best eight.

This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.

The feed-forward block (SwiGLU)

The feed-forward works per token

One expert: a narrower block

256 experts

The router: the top 8

Papers and sources