The perceptron: one neuron that learns
Start with a single neuron that learns to convert ancient Egyptian cubits to metres: a weight, a loss, backpropagation and gradient descent, which you can try by dragging the weight.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão . Open the page with scripts on to play the slides, one key press per idea.
The perceptron “In the beginning, there was the perceptron” A bare neuron: a line in, a line out. No numbers yet
One input, one output One thing in, one thing out Ask: what can one neuron do? How could it learn?
Data: cubits and metres Data: the same lengths in ancient Egyptian royal cubits and in metres Different tombs, rods, reigns → not perfectly proportional No exact answer; we want the best fit
What the neuron must learn We don’t know the conversion: let the computer learn it Cubits in, metres out
A weight: y = w · x A neuron can only multiply by a weight, w Guess: a cubit is a metre (truth ≈ 0.524) New counter: FLOPs per question = 1 multiplication
The forward pass One measurement: 2 cubits Blue = data flowing forward; the answer is 2.00 m
The error The data says 1.07 m (white dot from the table) Error = the gap: 0.93 m off
Loss: the squared error Loss = the error squared: 0.93² ≈ 0.86 Always positive; twice as far off = four times the loss This is the number to shrink
Backpropagation: the output Let the error flow backwards, one step at a time m = the answer; L = (m − t)² dL/dm = 2(m − t) = 1.86: the error signal
Backpropagation: the weight m = w·x, so dm/dw = x = 2 Chain rule: dL/dw = 1.86 × 2 = 3.72, the weight’s gradient Orange = signal flowing back Loss landscape: ball = current w, arrow points downhill
Your turn: drag the weight Nudging w against the gradient = gradient descent Best w ≈ 0.52; loss never hits zero (measurements disagree) Great Pyramid, 280 cubits, never seen: ≈ 147 m
Papers and sources Frank Rosenblatt (1958): The Perceptron: A Probabilistic Model for Information Storage and Organization in the Brain Ancient Egypt (c. 2700 BC): The royal cubit (about 0.52 m): the unit that built the pyramids Legendre · Gauss (1805 · 1809): The method of least squares Rumelhart, Hinton & Williams (1986): Learning representations by back-propagating errors Augustin-Louis Cauchy (1847): Méthode générale pour la résolution des systèmes d’équations simultanées (gradient descent)