ResNet: why deeper networks got worse, and the shortcut that fixed it
The degradation problem and vanishing gradients, measured on a small network, and how shortcut connections in ResNet give the residual stream that every transformer uses.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão . Open the page with scripts on to play the slides, one key press per idea.
Do more layers help? Do more layers work better? Our experiment: 64 neurons wide, 10,000 digits, vary the depth 2 layers ≈ 1.6 % train error; 4 layers half; 6 as good, test lower too Thick = train, dashed = test; average of 2 runs
Deeper gets worse Keep adding: 8, 12, 16 creep up; 32, 48, 64 shoot up It’s TRAINING error: not overfitting Deeper can’t learn what shallower did: the degradation problem (paper: 20 vs 56 layers)
The degradation problem Should be impossible: 20 trained layers + 36 identity layers = a 56-layer net just as good So deeper can match; the optimiser can’t find it Doing nothing is hard for weights + ReLUs
A plain stack Plain stack, no shortcuts; bars = gradient reaching each layer Last layer = reference (full signal) 8 layers: the first gets ≈ 1/10
Vanishing gradients: 16 layers 16 layers: the first gets ≈ 1/100 Each layer shrinks the gradient a little (Measured, period-typical initialisation)
Vanishing gradients: 32 layers 32 layers: the first gets ≈ 1 in 40,000 Early layers barely learn: vanishing gradient
Vanishing gradients: 56 layers 56 layers (the paper’s size): ≈ 1 in 100 million
Adding shortcuts Add shortcuts from the output end: the gradient returns to the early layers All in: first layer gets as much as the last (27 bypasses) y = x + f(x): layers learn only the change; the plus is a gradient highway Zero weights → block passes x through: a deep net can start shallow
The residual stream Shortcuts join into one line, input → output: the residual stream A shared notebook: each layer reads, adds a small correction Never replaced, only added to; the gradient runs straight along it This is how a transformer is drawn
The experiment again Experiment again with shortcuts: blue stays ≈ 1 % up to 64 layers Doesn’t improve with depth here, but no longer gets worse Paper: ResNet-34 beats 18; went on to 152 y = x + f(x) is in every transformer, DeepSeek’s too
Papers and sources He, Zhang, Ren & Sun (2015): Deep Residual Learning for Image Recognition (ResNet)