26: Convolutional Neural Networks
A note on this chapter
- An addition, and it fills the most conspicuous gap in the set
- The notes go from fully-connected networks (chapters 08-09) straight to transformers (chapter 21), skipping the architecture that started the deep learning era
- Chapter 20's ImageNet-era hardware and training toolkit exist because of this architecture
- The best way in is something these notes already built
- Chapter 18 detected pedestrians and text by cropping a patch, running a classifier, sliding the window, and repeating
- A convolutional network is that idea - with the detector learned, and the sliding built into the arithmetic
Why fully-connected fails on images
- The parameter explosion, in chapter 08's own numbers
- A 100 x 100 grayscale image is 10,000 inputs; chapter 08's quadratic-feature approach needed ~50,000,000 features
- A fully-connected layer is no better: 10,000 inputs to just 100 hidden units is a million weights in the first layer alone
- Overfitting (chapter 10) on anything but an enormous training set, and slow with it
- The deeper problem: a fully-connected net doesn't know it's looking at an image
- To chapters 08-09, pixel (3, 7) and pixel (90, 12) are just feature 307 and feature 9,012 - shuffle every image's pixels with the same permutation and the network learns exactly as well
- But images have structure: nearby pixels are related, and a cat in the top-left corner is the same cat as one in the centre
- A fully-connected net has to relearn every pattern at every position it might appear - the same edge detector, ten thousand times over
- Convolution builds both facts - locality and translation - into the architecture
- The same move as chapter 24's Invariant Point Attention: don't spend parameters learning something you can guarantee by construction
The convolution operation
- The ingredients
- A filter (or kernel) - a small grid of learned weights, typically 3 x 3 or 5 x 5
- Slide it over the image; at each position, take the element-wise product with the patch underneath and sum - one number per position
- The grid of those numbers is a feature map: where in the image the filter's pattern occurs
One output cell: the f × f filter W dotted with the image patch at position (i, j), plus a bias w0 — the course's usual subscript-0 convention. It is chapter 18's sliding window written as arithmetic — and each output cell is just one of chapter 08's logistic-unit sums, waiting for its activation function. (Strictly this is cross-correlation; the "convolution" of the name flips the filter first, and since W is learned the distinction changes nothing.)
- Two knobs
- Stride s - how far the window moves each step; stride 2 halves the output size
- Padding p - a border of zeros so the filter can centre on edge pixels; "same" padding keeps the output the size of the input
For an n × n input: a 6 × 6 image under a 3 × 3 filter with no padding and stride 1 gives (6 − 3)/1 + 1 = 4, a 4 × 4 map — checked by the code below.
import numpy as np
def conv2d(X, W, stride=1):
f = W.shape[0]
out = (X.shape[0] - f) // stride + 1
Y = np.zeros((out, out))
for i in range(out):
for j in range(out):
patch = X[i*stride : i*stride+f, j*stride : j*stride+f]
Y[i, j] = (patch * W).sum() # dot the filter with the patch
return Y
X = np.zeros((6, 6))
X[:, :3] = 1.0 # bright left half, dark right half
W = np.array([[1., 0, -1], # a vertical-edge filter:
[1., 0, -1], # bright-on-the-left minus
[1., 0, -1]]) # bright-on-the-right
conv2d(X, W) # [[0., 3., 3., 0.],
# [0., 3., 3., 0.],
# [0., 3., 3., 0.],
# [0., 3., 3., 0.]] - lights up exactly on the edge
Ten lines of chapter 03 matrix arithmetic. The filter responds with zero on the flat regions and 3 along the boundary — a feature detector. In a CNN this filter's nine numbers are not designed but learned by chapter 09's backpropagation, because "detect edges" is what reduces the classification cost.
Why it works: three ideas
- 1) Local connectivity
- Each output looks at an f × f patch, not the whole image - nearby pixels are the ones that form patterns together
- 2) Weight sharing
- The same nine numbers scan every position - one edge detector for the whole image, not ten thousand copies
- This is where the parameter count collapses
- 3) Translation equivariance
- Shift the input, and the feature map shifts with it - a pattern is detected wherever it appears, by construction, with no retraining
- The fully-connected net's relearn-it-everywhere problem, deleted
fc_layer = 10_000 * 100 + 100 # 1,000,100 - chapter 08's image into 100 FC units conv_layer = 3 * 3 * 1 * 16 + 16 # 160 - sixteen 3x3 filters cover the same image fc_layer // conv_layer # 6,250x fewer parameters
The weight-sharing arithmetic: a convolutional layer's parameter count depends on the filter size and the number of filters — not on the image size at all. Chapter 07's overfitting-vs-parameters trade-off, won by architecture instead of by penalty.
Pooling
- Between convolutions, shrink the maps
- Max pooling: split the map into (usually) 2 x 2 blocks and keep each block's maximum - "was the feature found anywhere in this neighbourhood?"
- Quarters the computation for the next layer, and makes the representation tolerant of small shifts - the feature's exact pixel matters less and less as you go deeper
- No parameters at all
def maxpool2(A): # 2x2 blocks, stride 2 return A.reshape(A.shape[0]//2, 2, A.shape[1]//2, 2).max(axis=(1, 3)) maxpool2(conv2d(X, W)) # [[3., 3.], # [3., 3.]] - "an edge was found in each quadrant"
The 4 × 4 edge map pooled to 2 × 2. Detail is deliberately thrown away — what survives is whether and roughly where the feature occurred, which is what the next layer needs.
The architecture
- The repeating unit: convolution → ReLU (chapter 20) → pool
- Each layer has many filters, so its output is a stack of feature maps - the channels. The next layer's filters see all channels at once: f × f × cin weights per filter
- As the network deepens, the maps get spatially smaller and the channels more numerous - less "where", more "what"
- The head is familiar
- Flatten the final maps into a vector, one or two fully-connected layers, then softmax over the classes - chapter 22's output head; chapter 08's one-hot classes
- The convolutional stages are a learned feature extractor; the head is chapter 06's classifier sitting on learned features - the arc from chapter 08 ("a network learns its own features") made literal
- Bookkeeping for a small digit classifier - chapter 18's character recognition problem, done properly
| Layer | Output shape | Parameters |
|---|---|---|
| input | 28 × 28 × 1 | 0 |
| conv 5×5, 8 filters + ReLU | 24 × 24 × 8 | 208 |
| max pool 2×2 | 12 × 12 × 8 | 0 |
| conv 5×5, 16 filters + ReLU | 8 × 8 × 16 | 3,216 |
| max pool 2×2 | 4 × 4 × 16 | 0 |
| flatten | 256 | 0 |
| fully connected, 100 + ReLU | 100 | 25,700 |
| softmax, 10 classes | 10 | 1,010 |
| total | 30,134 |
- For comparison: a single fully-connected layer from the raw 784 pixels to 100 units costs 78,500 parameters - more than this entire network, before it has classified anything
Training a CNN
- Nothing in this section is new, which is the point
- The loss is softmax cross-entropy (chapters 06 and 22); the gradients come from backpropagation (chapter 09) - weight sharing just means each filter weight accumulates gradient from every position it visited, exactly like the Δ accumulators of chapter 09
- The optimizer, initialization, normalization and dropout are chapter 20's toolkit; batch norm in particular grew up inside these networks
- Data augmentation - shifts, crops, flips, small distortions - is chapter 18's artificial data synthesis, and it works so well here precisely because the label survives those transformations
The ImageNet moment
- The pieces of this chapter are old - LeNet read cheques in the 1990s (chapter 08's "postcode reading" aside is this very architecture)
- What was missing was chapter 20's list: data, hardware, and the training toolkit
- 2012: AlexNet
- A CNN trained on GPUs, with ReLU and dropout, on ImageNet's million labelled images
- It nearly halved the best error rate in the field's flagship competition - not an increment, a discontinuity
- Chapter 11's "more data beats a cleverer algorithm" and chapter 20's toolkit, meeting the right architecture: this result is where the modern era of chapters 20-25 actually begins
- What the trained filters turned out to be
- First layer: edge and colour detectors - nobody asked for Sobel filters, but that is what gradient descent built, our hand-designed edge filter above rediscovered from data
- Deeper layers: textures, then parts, then whole objects - the feature hierarchy chapter 08 promised ("a network learns its own features"), visible in the weights
From sliding windows to convnets
- Chapter 18's pipeline, revisited with this chapter in hand
- The pedestrian detector cropped a patch, classified it, slid over, and repeated - re-computing everything for each of thousands of overlapping windows
- Convolution shares that computation: neighbouring windows overlap almost entirely, and a convolutional pass computes every window's features once. The whole sliding-window loop becomes a single forward pass
- Modern detection goes further
- One network predicts boxes and classes over the whole image directly - chapter 18's separate detect / segment / classify stages collapse into one model trained end-to-end
- Ceiling analysis (chapter 18) still applies, but the ceiling being analysed is now one network's components rather than a pipeline of hand-built stages
CNNs and transformers
- The two architectures make opposite bets
- A convolution hard-codes which positions interact (a local window, everywhere the same); attention (chapter 21) learns which positions interact, per input
- Convolution's assumptions - locality, translation - are an inductive bias: chapter 10's bias/variance trade-off applied to architecture. Strong assumptions help when data is scarce, and cost you when data is abundant and the assumptions bind
- Vision transformers
- Cut the image into 16 × 16 patches, embed each patch as a token, and run chapter 21's transformer on the sequence - images entering by the front door of chapter 22's pipeline
- With enough data they match or beat CNNs; with little data the CNN's built-in bias still wins - exactly what the chapter 10 framing predicts
- In practice the bet is often hedged: convolutional early layers for cheap local features, attention above them - and multimodal assistants (chapter 20) ingest images through exactly such vision encoders
Summary
- Fully-connected layers scale with image size and relearn every pattern at every position; convolution fixes both by construction
- A convolution slides a small learned filter everywhere - chapter 18's sliding window as arithmetic, with the detector learned by chapter 09's backprop
- Local connectivity, weight sharing and translation equivariance are the three ideas; the parameter count drops thousands-fold (6,250× in the worked example)
- Pooling throws away exact position to keep "was it there" - tolerance and cheapness at once
- Stack conv-ReLU-pool, finish with softmax: a learned feature hierarchy under chapter 06's classifier, trained with chapter 20's toolkit and chapter 18's augmentation
- AlexNet on ImageNet in 2012 is where this architecture met enough data and compute - the opening event of everything in chapters 20-25
- Convolution is a hard-coded, local, everywhere-identical attention pattern; the transformer learns its pattern instead. Which wins is a data-budget question, and chapter 10 already gave the framework for it