holehouse.org Blog Machine learning notes

26: Convolutional Neural Networks

A note on this chapter

Why fully-connected fails on images

The convolution operation

Yij=∑a=1f∑b=1fWabXi+a,j+b+w0 One output cell: the f × f filter W dotted with the image patch at position (i, j), plus a bias w0 — the course's usual subscript-0 convention. It is chapter 18's sliding window written as arithmetic — and each output cell is just one of chapter 08's logistic-unit sums, waiting for its activation function. (Strictly this is cross-correlation; the "convolution" of the name flips the filter first, and since W is learned the distinction changes nothing.)
output size=⌊n+2p−fs⌋+1 For an n × n input: a 6 × 6 image under a 3 × 3 filter with no padding and stride 1 gives (6 − 3)/1 + 1 = 4, a 4 × 4 map — checked by the code below.
import numpy as np

def conv2d(X, W, stride=1):
    f = W.shape[0]
    out = (X.shape[0] - f) // stride + 1
    Y = np.zeros((out, out))
    for i in range(out):
        for j in range(out):
            patch = X[i*stride : i*stride+f, j*stride : j*stride+f]
            Y[i, j] = (patch * W).sum()   # dot the filter with the patch
    return Y

X = np.zeros((6, 6))
X[:, :3] = 1.0                  # bright left half, dark right half

W = np.array([[1., 0, -1],      # a vertical-edge filter:
              [1., 0, -1],      # bright-on-the-left minus
              [1., 0, -1]])     # bright-on-the-right

conv2d(X, W)   # [[0., 3., 3., 0.],
               #  [0., 3., 3., 0.],
               #  [0., 3., 3., 0.],
               #  [0., 3., 3., 0.]]  - lights up exactly on the edge

Ten lines of chapter 03 matrix arithmetic. The filter responds with zero on the flat regions and 3 along the boundary — a feature detector. In a CNN this filter's nine numbers are not designed but learned by chapter 09's backpropagation, because "detect edges" is what reduces the classification cost.

Why it works: three ideas

fc_layer   = 10_000 * 100 + 100     # 1,000,100 - chapter 08's image into 100 FC units
conv_layer = 3 * 3 * 1 * 16 + 16    # 160       - sixteen 3x3 filters cover the same image

fc_layer // conv_layer              # 6,250x fewer parameters

The weight-sharing arithmetic: a convolutional layer's parameter count depends on the filter size and the number of filters — not on the image size at all. Chapter 07's overfitting-vs-parameters trade-off, won by architecture instead of by penalty.

Pooling

def maxpool2(A):                    # 2x2 blocks, stride 2
    return A.reshape(A.shape[0]//2, 2, A.shape[1]//2, 2).max(axis=(1, 3))

maxpool2(conv2d(X, W))   # [[3., 3.],
                         #  [3., 3.]]  - "an edge was found in each quadrant"

The 4 × 4 edge map pooled to 2 × 2. Detail is deliberately thrown away — what survives is whether and roughly where the feature occurred, which is what the next layer needs.

The architecture

image conv maps pool conv pool flatten fully connected + softmax "cat" edges → textures → parts → objects: each stage detects patterns in the previous stage's detections
Convolution and pooling repeated, then a small classifier on top. Spatial size falls, channel count grows — the network trades "where" for "what".
A LeNet-style network for 28 × 28 digit images. Almost all parameters sit in the little head; the feature extractor is nearly free.
LayerOutput shapeParameters
input28 × 28 × 10
conv 5×5, 8 filters + ReLU24 × 24 × 8208
max pool 2×212 × 12 × 80
conv 5×5, 16 filters + ReLU8 × 8 × 163,216
max pool 2×24 × 4 × 160
flatten2560
fully connected, 100 + ReLU10025,700
softmax, 10 classes101,010
total30,134

Training a CNN

The ImageNet moment

From sliding windows to convnets

CNNs and transformers

Summary