ResNet: why deeper networks got worse, and the shortcut that fixed it

The degradation problem and vanishing gradients, measured on a small network, and how shortcut connections in ResNet give the residual stream that every transformer uses.

This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.

Do more layers help?

Deeper gets worse

The degradation problem

A plain stack

Vanishing gradients: 16 layers

Vanishing gradients: 32 layers

Vanishing gradients: 56 layers

Adding shortcuts

The residual stream

The experiment again

Papers and sources