Why a chain of neurons adds nothing, and a bend does

Real temperature data from Norwich shows why stacking straight-line neurons never helps, and how a ReLU bend in every neuron lets a network fit a curve.

This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão. Open the page with scripts on to play the slides, one key press per idea.

Data: Norwich temperatures

One neuron, and a bias

Two neurons in a chain

A wider layer: still a line

The bend: ReLU

Training the layer

Papers and sources