20: Modern Deep Learning
A note on this chapter
- Chapters 01 to 19 are my write-up of Professor Ng's 2011 course - this one isn't
- It's a high-level tour of what happened next
- The course predates essentially all of it
- But almost none of it is conceptually new - it's the same machinery at a scale nobody had tried yet
- Deliberately shallow
- Enough to know what each thing is, roughly how it works, and what it's for
- The details live in their own chapters now: attention (21), the mathematics of LLMs (22), diffusion models (23), AlphaFold2 (24) and protein language models (25)
- This chapter is the map - plus the one set of details that belongs here: the training toolkit that separates 2011's networks from deep ones
- I've avoided version numbers throughout
- They date within months, and the concepts don't
What actually changed
- Four things arrived together, and none of them alone would have been enough
- Scale - models with billions to trillions of parameters, trained on a large fraction of the public internet
- Hardware - GPUs, then accelerators built specifically for this, making the matrix multiplies of chapter 03 fast enough to be practical at that size
- Architecture - the transformer (chapter 21), which unlike the recurrent models before it parallelizes across the whole sequence
- Training know-how - the numerical fixes that keep signals and gradients alive at depth: new activations, matched initialization, adaptive optimizers, normalization. The five sections below cover these properly - they are the bridge between chapter 09's networks and everything after
- The consequence that mattered most was a change in how you use a model
- Old way - collect labelled data for your task, train a model for that task
- New way - take a model someone else trained on an enormous unlabelled corpus, and adapt it
- This is transfer learning, and it's why a lab with no ML budget can now do serious ML
- Where the labels went
- Labelled data was always the bottleneck - chapters 10 and 11 are largely about spending a scarce labelling budget well
- The trick is to invent a label the data already contains: hide part of the input, predict it from the rest
- Next word, masked word, noised image - the supervision comes free with the data
- Usually called self-supervised learning, and it's the single idea underneath everything below
Activations: from sigmoid to ReLU
- The notes' networks (chapters 08 and 09) use sigmoid units throughout - and that is a large part of why they could not go deep
- Backpropagation multiplies one derivative factor per layer (chapter 09)
- The sigmoid's derivative is g′(z) = a(1 − a) - the identity chapter 09 verified numerically - and it is at most 1/4, at z = 0
- Away from zero it saturates and the derivative is nearly nothing - the same flat-region problem chapter 21 met in softmax
import numpy as np 0.25 ** 30 # 8.7e-19 - the best-case gradient factor through 30 sigmoid layers 1.0 ** 30 # 1.0 - the same product for ReLU units on their active side
The whole story in two lines: sigmoid depth multiplies the gradient into oblivion; ReLU depth leaves it alone.
- The fix is almost embarrassingly blunt: the rectified linear unit
- g(z) = max(0, z) - identity for positive input, zero otherwise
- Derivative is exactly 1 on the active side, exactly 0 on the inactive side - no saturation, no shrinking factor, and cheaper than an exponential
- Not differentiable at exactly 0, which in practice matters not at all (pick either side)
- Its one failure mode, and the descendants
- A unit whose input is always negative outputs 0 forever and its gradient is 0 forever - a dead unit
- Leaky ReLU gives the negative side a small slope so nothing can die; GELU (a smooth blend of the two regimes) is the default inside transformers (chapters 21-22)
- All of them keep the property that matters: a derivative near 1 over most of the working range
Initialization: keeping the signal alive
- Chapter 09 said: initialize to "small random values" - random for symmetry-breaking, which stands. But how small becomes critical with depth
- Each layer multiplies the signal's variance by (number of inputs) × Var(w) - the same sum-of-independent-terms argument chapter 21 used to justify dividing by √dk
- If that factor is below 1 the activations shrink geometrically; above 1 they explode geometrically. Forty layers turn "slightly off" into "gone"
rng = np.random.default_rng(0) n = 256 # 256 units per layer def depth_test(scale, act): x = rng.standard_normal(n) for _ in range(40): # forty layers deep W = rng.standard_normal((n, n)) * scale x = act(W @ x) return x.std() relu = lambda z: np.maximum(0, z) depth_test(0.01, relu) # 8e-39 - "small values": the signal is gone depth_test(np.sqrt(2 / n), relu) # 0.27 - He init: still alive at layer 40 depth_test(3 * np.sqrt(2 / n), relu) # 6e+18 - too big: exploded instead
Chapter 09's advice, stress-tested at depth. The window between vanishing and exploding is narrow, and the He formula puts you in it for any width automatically.
Optimizers: momentum, RMSProp, Adam
- The problem, which these notes have already drawn
- Chapter 04's contour plot: badly scaled features make long thin valleys, and gradient descent zig-zags across the steep direction while crawling along the shallow one
- The learning rate is capped by the steepest direction, so the shallow direction sets the runtime
- Feature scaling fixed this for the inputs; inside a deep network the same mismatch reappears at every layer, where you cannot hand-scale it away
- Momentum - remember the direction you have been moving
- Keep a running velocity; the zig-zag components cancel in the average, the consistent component accumulates
- RMSProp - give every parameter its own learning rate
- Track a running average of each parameter's squared gradient and divide by its square root - parameters with habitually large gradients get small steps, and vice versa
- Adam - both at once, plus a correction
- A momentum-style average m of the gradient, an RMSProp-style average v of its square, and a correction for the early steps when both averages are still warming up from zero
def race(update, L, tol=1e-8, steps=100_000):
theta, state = np.array([10.0, 1.0]), {}
for k in range(1, steps + 1): # J = theta1^2/2 + L*theta2^2/2
grad = np.array([theta[0], L * theta[1]])
theta = update(theta, grad, state, k)
if theta[0]**2 / 2 + L * theta[1]**2 / 2 < tol:
return k
def gd(t, g, s, k): # best stable learning rate
return t - (2 / (L + 1)) * g
def momentum(t, g, s, k):
s['v'] = (4/9) * s.get('v', 0) + g # textbook-optimal beta for L = 25
return t - (1/9) * s['v']
def adam(t, g, s, k):
s['m'] = 0.9 * s.get('m', 0) + 0.1 * g
s['v'] = 0.999 * s.get('v', 0) + 0.001 * g * g
m_hat = s['m'] / (1 - 0.9 ** k)
v_hat = s['v'] / (1 - 0.999 ** k)
return t - 0.5 * m_hat / (np.sqrt(v_hat) + 1e-8)
L = 25 # a mildly elongated bowl - chapter 04's contours
race(gd, 25) # 141 steps
race(momentum, 25) # 37 steps - the sqrt(condition-number) speedup
race(adam, 25) # 172 steps - unremarkable here...
L = 10_000 # ...but make the scales wildly different and
race(gd, 10_000) # 67,370 steps - capped by the steep axis
race(adam, 10_000) # 220 steps - per-parameter scaling wins
An honest race. On a clean, mildly ill-conditioned bowl, well-tuned momentum is the fastest thing there is and Adam is nothing special. Blow the conditioning out to 104 — the un-scaled-features regime of chapter 04, which deep networks recreate internally — and Adam wins by 300× without retuning anything. Robustness, not raw speed, is why it became the default.
Learning-rate schedules and warmup
- Chapter 02 argued no schedule is needed: near a minimum the gradient shrinks, so steps shrink themselves
- True for batch descent on a smooth bowl. With mini-batches (chapter 17) it fails: gradient noise does not shrink as you converge, so a constant rate leaves you rattling around the minimum - chapter 17 already met this as SGD "wandering"
- Decay - end low
- Chapter 17's fix was α = c1/(t + c2); the modern default is cosine decay - a smooth run from the peak rate down to near zero over the planned training length
- Same idea either way: big steps to cross the landscape early, small steps to settle late
- Warmup - start low too
- The first few hundred steps ramp the rate up from zero
- Early on, Adam's running averages are built from a handful of noisy mini-batches, and a full-size step taken on garbage statistics can wreck the network before training starts - warmup lets the estimates settle first
- So the standard schedule is a ramp up then a long cosine down - and its length is a hyperparameter chosen on validation data, exactly the chapter 10 procedure
Dropout and batch normalization
- Dropout - regularization by sabotage
- During training, independently zero each hidden unit with probability 1 − p (keep with probability p), a fresh coin flip every example
- No unit can rely on a specific partner existing, so the network cannot build the brittle co-adapted features that overfitting (chapters 07 and 10) is made of
- Equivalent view: you are training a huge ensemble of thinned networks that share weights, and averaging them at test time
a = np.ones(100_000) * 2.0 # an activation of 2.0, many trials p = 0.8 # keep 80% of units mask = rng.random(100_000) < p (a * mask / p).mean() # 2.0 - the expectation survives
The scaling checked: dropping a fifth of the units while dividing by 0.8 leaves the expected signal exactly where it was.
- Batch normalization - feature scaling, moved inside the network
- Chapter 04 standardized the inputs so gradient descent saw round contours. A deep network un-does that favour internally: each layer's inputs are the previous layer's outputs, with whatever mean and scale training has drifted them to
- So re-standardize between layers, using the current mini-batch's own statistics
- What it buys
- Much higher usable learning rates and far less sensitivity to initialization - training that simply works where it used to diverge
- A mild regularization side-effect, since each example's normalization depends on which batch-mates it drew
- Its sibling
- Layer normalization (chapter 21) computes the same statistics per example across features instead of per feature across the batch - no batch dependence, which is why transformers use it
- Batch norm rules convolutional vision models; layer norm rules sequence models. Same equation, different axis
- The recipe shift, side by side
| These notes, 2011 | Modern practice | |
|---|---|---|
| Activation | sigmoid / tanh | ReLU family (GELU in transformers) |
| Initialization | "small random values" | He / Xavier - variance matched to width |
| Optimizer | batch gradient descent, fminunc | Adam (or SGD + momentum) on mini-batches |
| Learning rate | one constant α, chosen by plot | warmup then cosine decay |
| Regularization | L2 penalty (chapter 07) | weight decay + dropout + early stopping + data augmentation (chapter 18) |
| Normalization | inputs only (chapter 04) | batch/layer norm between layers |
| Feasible depth | a few layers | hundreds |
- Nothing in this section is conceptually deep, and that is the point
- The 2011 course had the right objective, the right algorithm and the right architecture idea; what was missing was a handful of numerical fixes that keep signals and gradients alive at depth
- Those fixes, plus the hardware and data above, are the gap between chapter 09's networks and everything in chapters 21-25
Large language models
This section is the overview; chapter 22 does the mathematics properly.
- The training objective is almost insultingly simple
- Given a run of text, predict the next token
- A token is roughly a word-piece - common words are one token, rare ones split into several
- That's it. Softmax over the vocabulary, cross-entropy loss - the multi-class classification of chapter 09, with a vocabulary of maybe 100,000 classes
- Why that produces something so much more general than it sounds
- To predict the next token well across the whole internet, you have to model whatever generated it
- Finishing "the capital of France is" needs a fact; finishing a proof needs the argument; finishing a function body needs the code to typecheck
- So syntax, facts, reasoning patterns and style all fall out of one objective, because all of them reduce prediction error
- Scaling laws - the empirical finding that drove the whole build-out
- Loss falls predictably as a power law in model size, data and compute
- Predictably enough to plan a training run before doing it, which is what made the capital expenditure defensible
- Note this is an empirical regularity over many orders of magnitude, not a theorem - it holds until it doesn't
- In-context learning - the surprise
- Put a few worked examples in the prompt and the model does the task, with no gradient step and no weight change
- Nobody trained for this - it emerged from scale
- Practically it means the interface to the model is text, not a training pipeline
What they're used for
- Writing, editing, summarizing, translating
- Code - generation, review, refactoring, explanation; probably the strongest commercial use
- Extraction - pulling structure out of unstructured text, which used to be a bespoke NLP project each time
- Classification with no training set, by simply describing the classes
- An interface layer - natural language over an API, a database, or a pile of documents
How a raw model becomes an assistant
- A model straight out of pretraining is not a chatbot
- It continues text. Ask it a question and a plausible continuation is a list of similar questions
- It's a model of the corpus, not an assistant - being helpful was never the objective
- So there's a second stage, usually called post-training
- Supervised fine-tuning
- Continue training on curated examples of instructions and good responses
- Ordinary supervised learning - the model learns the shape of being asked and answering
- Learning from preferences
- Show humans two candidate responses, ask which is better
- Train a reward model to predict that judgement, then optimize the model against it - this is RLHF
- Why preferences rather than labels: "which of these is better" is a question people can answer reliably, "write the ideal response" is not
- Reasoning training
- More recent, and the reason for the recent step-change on maths and code
- Reward the model for reaching a verifiably correct answer, letting it work at length first
- The model learns to spend more computation on harder problems - test-time compute becomes a dial you can turn
- Supervised fine-tuning
- Alignment is the open problem here
- The objective is a proxy for what we want, and optimizing hard against a proxy is how you get a model that games it
- Same failure mode as any badly-chosen cost function, with more consequences
The assistants - ChatGPT, Claude, Gemini
- All three are the same recipe - a large transformer, pretrained on text, then post-trained into an assistant
- They differ in training data, in post-training method and emphasis, and in the surrounding product
- Not in any deep architectural sense that would matter to this chapter
- ChatGPT (OpenAI)
- The one that made this public in late 2022, built on the GPT model family
- Its real contribution was the interface - the underlying capability existed before, but nobody had put a chat box on it
- Claude (Anthropic)
- Model families named Opus, Sonnet and Haiku - roughly most capable, balanced, and fastest
- Notable for Constitutional AI - the model critiques and revises its own outputs against an explicit written set of principles, so some of the human feedback loop is replaced by AI feedback against a stated standard
- Strong on long-context work and on code
- Gemini (Google DeepMind)
- Natively multimodal - text, images, audio and video handled by one model rather than bolted together
- Long context windows, and tight integration with Google's own products
- Two things worth understanding about all of them
- Context window - how much text the model can attend to at once. This is the practical constraint you feel most often, and it's bounded by the O(n2) cost of attention (chapter 21)
- Multimodality - images, audio and video get turned into token sequences too, so the same machinery applies. Nothing conceptually new; the pipeline just has more front ends
Agentic workflows
- The shift from "model that answers" to "model that does"
- Give the model a set of tools - run code, search, read a file, call an API
- It emits a structured call, your code executes it, the result goes back into the context, and it continues
- Loop until the task is done
- Why this is more than a convenience
- It closes the loop with reality - the model can now check its own work rather than assert it
- It fixes the things a language model is inherently bad at, by delegating them: arithmetic to a calculator, current facts to a search, correctness to a test suite
- And it makes the work incremental - errors surface at step three rather than at the end
- The pieces you'll hear named
- Tool use / function calling - the model outputs a structured call rather than prose
- RAG (retrieval-augmented generation) - fetch relevant documents, put them in the context, answer from them. Grounds answers in a source you control, and gets around the context window for large corpora
- MCP (Model Context Protocol) - an open standard for how tools and data sources describe themselves to models, so integrations aren't rebuilt per vendor
- Sub-agents - delegating a self-contained piece of work to a separate context, which keeps the main one small
- What they're used for
- Coding agents that read a repository, make changes and run the tests
- Research - search, read, cross-check, synthesize
- Data work - write the query, run it, plot the result, notice it looks wrong, fix it
- Anything that was a script you'd have written by hand, where the steps aren't known in advance
- The failure mode to design around
- Errors compound - a 95% reliable step is 60% reliable after ten of them
- So the engineering is mostly about verification, recovery and keeping the human in the loop where it matters, not about the prompt
What to be sceptical about
- Evaluation is much harder than it looks
- Chapter 10's discipline - train/validation/test, held out properly - matters more than ever, and is followed less
- When the training set is "the internet", your test set is probably in it. Benchmark contamination is pervasive and often undetectable
- In biology the leak is subtler: random splits of sequence data put homologues on both sides, so you measure memorization and call it generalization
- Fluency is not correctness
- These models are optimized to produce plausible continuations, and a confident wrong answer is exactly as fluent as a right one
- There's no internal signal that separates the two for you - hence tools, retrieval and verification
- Bias in, bias out - unchanged since chapter 11, just at larger scale and harder to inspect
- Cost and access
- Pretraining a frontier model is out of reach for almost everyone
- But using one is cheap, and fine-tuning an open one is very achievable - which is the practically important fact
Summary
- One idea underlies nearly all of it
- Invent a supervised task the unlabelled data answers for itself, train an enormous model on it, then adapt
- Next token for text, masked residue for proteins, added noise for images
- LLMs - next-token prediction at scale, post-trained into assistants; used for text, code, extraction, and as a natural-language interface to everything else
- Agentic workflows - give the model tools and a loop, so it can act and check rather than only answer
- Diffusion - learn to remove a little noise, run it backwards to generate; now chapter 23, with a working NumPy model
- AlphaFold2 - sequence to structure via evolutionary covariation; now chapter 24, piece by piece
- ESM - masked language modelling on protein sequences; now chapter 25, ESM2 and ESM-C compared
- What carries over from the rest of these notes
- Everything. Gradient descent, regularization, bias and variance, the train/validation/test discipline, feature scaling, softmax, evaluation metrics for skewed data
- The models got much bigger; the failure modes are the ones in chapters 07, 10 and 11
- The one genuinely new mechanism is attention (chapter 21) - and chapters 22-25 show how far it travels