holehouse.org Blog Machine learning notes

20: Modern Deep Learning

A note on this chapter

What actually changed

Activations: from sigmoid to ReLU

g′(z)=g(z)(1−g(z))≤14 The best case for a sigmoid layer. Chain 30 of them and the gradient reaching the early layers carries a factor of at most (1/4)30 — the vanishing gradient problem. The early layers stop learning not because of a bug but because of arithmetic.
import numpy as np

0.25 ** 30   # 8.7e-19 - the best-case gradient factor through 30 sigmoid layers
1.0  ** 30   # 1.0     - the same product for ReLU units on their active side

The whole story in two lines: sigmoid depth multiplies the gradient into oblivion; ReLU depth leaves it alone.

Initialization: keeping the signal alive

Var(w)=2nin He initialization for ReLU layers: variance 2 over the number of inputs, the 2 compensating for ReLU zeroing half its inputs. (Xavier initialization, 1/nin, is the same argument for symmetric activations.) Chosen so the layer's output variance equals its input variance — the multiply-per-layer factor is pinned at 1 by construction.
rng = np.random.default_rng(0)
n = 256                                     # 256 units per layer

def depth_test(scale, act):
    x = rng.standard_normal(n)
    for _ in range(40):                     # forty layers deep
        W = rng.standard_normal((n, n)) * scale
        x = act(W @ x)
    return x.std()

relu = lambda z: np.maximum(0, z)

depth_test(0.01,              relu)   # 8e-39  - "small values": the signal is gone
depth_test(np.sqrt(2 / n),    relu)   # 0.27   - He init: still alive at layer 40
depth_test(3 * np.sqrt(2 / n), relu)  # 6e+18  - too big: exploded instead

Chapter 09's advice, stress-tested at depth. The window between vanishing and exploding is narrow, and the He formula puts you in it for any width automatically.

Optimizers: momentum, RMSProp, Adam

v:=βv+∇J(θ),θ:=θ−αv β ≈ 0.9: each step is mostly the previous step plus a gradient nudge — a heavy ball rolling downhill rather than a walker re-deciding direction from scratch.
s:=ρs+(1−ρ)(∇J)2,θ:=θ−αs+ε∇J All operations element-wise. This is feature scaling (chapter 04) applied to the gradient, continuously, per parameter — the fix moved from preprocessing into the optimizer itself.
m:=β1m+(1−β1)∇J,v:=β2v+(1−β2)(∇J)2 m^=m1−β1t,v^=v1−β2t,θ:=θ−αm^v^+ε Defaults β1 = 0.9, β2 = 0.999 barely ever change — which is the actual selling point: it works out of the box across wildly different problems, at the cost of one extra stored value per parameter for each of m and v. This is the optimizer behind chapters 22-25.
def race(update, L, tol=1e-8, steps=100_000):
    theta, state = np.array([10.0, 1.0]), {}
    for k in range(1, steps + 1):          # J = theta1^2/2 + L*theta2^2/2
        grad = np.array([theta[0], L * theta[1]])
        theta = update(theta, grad, state, k)
        if theta[0]**2 / 2 + L * theta[1]**2 / 2 < tol:
            return k

def gd(t, g, s, k):                        # best stable learning rate
    return t - (2 / (L + 1)) * g

def momentum(t, g, s, k):
    s['v'] = (4/9) * s.get('v', 0) + g     # textbook-optimal beta for L = 25
    return t - (1/9) * s['v']

def adam(t, g, s, k):
    s['m'] = 0.9   * s.get('m', 0) + 0.1   * g
    s['v'] = 0.999 * s.get('v', 0) + 0.001 * g * g
    m_hat = s['m'] / (1 - 0.9 ** k)
    v_hat = s['v'] / (1 - 0.999 ** k)
    return t - 0.5 * m_hat / (np.sqrt(v_hat) + 1e-8)

L = 25                       # a mildly elongated bowl - chapter 04's contours
race(gd, 25)                 # 141 steps
race(momentum, 25)           #  37 steps - the sqrt(condition-number) speedup
race(adam, 25)               # 172 steps - unremarkable here...

L = 10_000                   # ...but make the scales wildly different and
race(gd, 10_000)             # 67,370 steps - capped by the steep axis
race(adam, 10_000)           #    220 steps - per-parameter scaling wins

An honest race. On a clean, mildly ill-conditioned bowl, well-tuned momentum is the fastest thing there is and Adam is nothing special. Blow the conditioning out to 104 — the un-scaled-features regime of chapter 04, which deep networks recreate internally — and Adam wins by 300× without retuning anything. Robustness, not raw speed, is why it became the default.

Learning-rate schedules and warmup

Dropout and batch normalization

a:=a⊙maskp Inverted dropout: divide by the keep-probability during training so the expected activation is unchanged — then test time needs no adjustment at all, just switch the mask off.
a = np.ones(100_000) * 2.0            # an activation of 2.0, many trials
p = 0.8                               # keep 80% of units
mask = rng.random(100_000) < p
(a * mask / p).mean()                 # 2.0 - the expectation survives

The scaling checked: dropping a fifth of the units while dividing by 0.8 leaves the expected signal exactly where it was.

x^=x−μBσB2+ε,y=γx^+β Chapter 04's mean-normalization formula, verbatim — μB and σB are the mini-batch's mean and standard deviation — plus a learned scale γ and shift β so the network can undo the normalization wherever it turns out to be unhelpful. At test time the batch statistics are replaced by running averages kept during training.
The 2011 recipe (chapters 06-09) against the modern one — every row is a small fix, and depth needs all of them at once.
These notes, 2011Modern practice
Activationsigmoid / tanhReLU family (GELU in transformers)
Initialization"small random values"He / Xavier - variance matched to width
Optimizerbatch gradient descent, fminuncAdam (or SGD + momentum) on mini-batches
Learning rateone constant α, chosen by plotwarmup then cosine decay
RegularizationL2 penalty (chapter 07)weight decay + dropout + early stopping + data augmentation (chapter 18)
Normalizationinputs only (chapter 04)batch/layer norm between layers
Feasible deptha few layershundreds

Large language models

This section is the overview; chapter 22 does the mathematics properly.

What they're used for

How a raw model becomes an assistant

internet- scale text pretraining predict next token fine-tuning curated examples preferences / reasoning assistant months, enormous cost comparatively cheap - this is where behaviour is shaped
Almost all the capability comes from stage one; almost all the behaviour from the stages after it.

The assistants - ChatGPT, Claude, Gemini

Agentic workflows

task model picks a tool run code / search / read result returns to the context done? answer
The loop - and every pass through it is a chance to catch an error, or to compound one.

What to be sceptical about

Summary