AlexNet: the same recipe at ImageNet scale
AlexNet (2012) is the digit reader scaled up to a thousand kinds of photograph; here a pre-trained image network runs live in the browser on test pictures or your camera.
This is the text of a chapter of the animated talk A Deep Seek into LLM Architecture by Ruben Galvão . Open the page with scripts on to play the slides, one key press per idea.
AlexNet (2012) 2012, built for ImageNet: 1.2 M photos, 1,000 classes Same recipe: 227×227×3 in; conv, ReLU, pool; channels 96, 256, 384, 384, 256 Flatten 9,216 → 4,096 → 4,096 → 1,000 + softmax ≈ 60 M weights (10,000× the digit reader); ≈ 1.4 G ops per image (≈ 3,650×) Maps after the first are illustrative
The real weights The real AlexNet (61 M weights) running in this page, offline Centre crop only (white frame); top-1 ≈ 55–57 %, so we show top five Hosted version runs MobileNetV2 (≈ 3.5 M) instead 2012: top-5 error 15.3 % vs 26.2 % for the runner-up: deep learning takes over vision Next: if depth helped, add more
Papers and sources Deng et al. (2009): ImageNet: A large-scale hierarchical image database Krizhevsky, Sutskever & Hinton (2012): ImageNet Classification with Deep Convolutional Neural Networks (AlexNet) BVLC · ONNX Model Zoo (pre-trained weights): The Caffe reference AlexNet, converted to ONNX