Why a chain of neurons adds nothing, and a bend does
Real temperature data from Norwich shows why stacking straight-line neurons never helps, and how a ReLU bend in every neuron lets a network fit a curve.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão . Open the page with scripts on to play the slides, one key press per idea.
Data: Norwich temperatures Real data: monthly average temperature, Norwich, ten years One dot per month per year; the thermometer replays a year Cold → warm → cold, but with scatter: no curve hits every dot Best we can do: a good typical guess
One neuron, and a bias One neuron; inputs: the month and a constant 1 (the bias) m = w·month + b·1; the bias lets the line sit at the right height Train: error, gradient back, nudge; the line settles Best straight line still misses by ≈ 4.5° Cost: 2 operations. Maths below: x → W → answer; violet dashed = learned weights
Two neurons in a chain Second neuron after the first: each column of neurons = a layer It has no bias of its own: the first one has it Matrices: W₁ makes the hidden h, w₂ makes the answer Nothing improves: two weights multiplied = one weight; a chain of lines is a line
A wider layer: still a line Make the layer wider, one neuron at a time Each new neuron: same input and bias; W₁ gains a row Counter: 3 weights per neuron, 24 for eight Line doesn’t move: a sum of lines is a line. Width alone doesn’t help
The bend: ReLU Fix: a bend in every neuron: ReLU (negative → 0) Icon on the line from W₁ to h; one comparison per neuron From the digit reader on, count only multiplies and adds New neurons nearly silent yet, but each can bend in its own place
Training the layer Train: error flows back, hinges move, the sum bends around the year Miss drops from ≈ 4.5° to ≈ 1°; the floor is the yearly scatter Bends added together make almost any curve: the whole trick Cost: 39 operations vs 2 (31 multiplies/adds + 8 bends)
Papers and sources Minsky & Papert (1969): Perceptrons: what linear units cannot do Open-Meteo (2015–2024): Historical daily temperatures at Norwich (ERA5 reanalysis) Nair & Hinton (2010): Rectified Linear Units Improve Restricted Boltzmann Machines George Cybenko (1989): Approximation by superpositions of a sigmoidal function