23: Diffusion Models
A note on this chapter
- An addition, like chapters 20-22 - the course predates all of this
- Chapter 20 used to hold a short survey of diffusion; this chapter replaces it with the actual mathematics and a working model
- Everything needed is already in these notes
- Gaussians from chapter 15, squared-error cost from chapter 02, a small neural network from chapter 08 trained with chapter 09's backpropagation, mini-batch descent from chapter 17
- The NumPy model below is genuinely nothing but those pieces
The idea
- The dominant approach to generating images, video and - increasingly - molecular structures
- Two processes, one fixed and one learned
- Forward process - take a real example and add a little Gaussian noise. Repeat a few hundred times. You end up with pure noise, and it's destroyed the same way every time
- Reverse process - train a network to undo one step of that: given a noisy example, predict the noise that was added
- Then generate by starting from pure noise and running the learned reverse step repeatedly, until a sample falls out
- Why go the long way round
- Going from noise to a photograph in one jump is a very hard function to learn
- Going from slightly-noisier to slightly-less-noisy is an easy one - and you can compose it as many times as you like
- The generation is decomposed into many small, individually-easy steps. That's the whole trick
The forward process
- Fix a noise schedule - a sequence of small variances β1, …, βT
- Each step shrinks the signal slightly and adds a little Gaussian noise
- The property that makes training practical
- A chain of Gaussians is a Gaussian, so you can jump straight from x0 to any step t in closed form - no need to simulate the chain
Learning to reverse it
- The network's one job: look at a noisy example and the step number, and predict the noise that was mixed in
- Call it εθ(xt, t)
- Training data is free in unlimited quantities: take a real x0, pick a random t, draw ε, mix them with the closed form above - now you know the right answer exactly
- Self-supervision again (chapter 20): the label is manufactured from the data itself
- Generation runs the chain backwards
- Start from xT ∼ N(0, I), and at each step subtract out (a correctly-scaled portion of) the predicted noise, then add back a little fresh noise
A complete diffusion model in NumPy
- The pieces above, assembled and run
- Data: one-dimensional, a mixture of two Gaussians - half the points near −2, half near +2
- Deliberately chosen so success is checkable: samples from a trained model must come out bimodal, and nothing simpler than a real generative model produces that from pure noise
- First, the data and the forward process
import numpy as np rng = np.random.default_rng(0) def sample_data(n): # half near -2, half near +2 modes = rng.choice([-2.0, 2.0], size=n) return modes + 0.3 * rng.standard_normal(n) T = 50 beta = np.linspace(1e-3, 0.2, T) # the noise schedule alpha = 1 - beta alpha_bar = np.cumprod(alpha) # alpha_bar[-1] = 0.0045: 0.45% of signal left x0 = sample_data(10000) eps = rng.standard_normal(10000) xT = np.sqrt(alpha_bar[-1]) * x0 + np.sqrt(1 - alpha_bar[-1]) * eps xT.mean(), xT.std() # (0.001, 1.0) - the data is gone; # every x0 ends as standard normal noise
The closed-form jump verified: after 50 steps the bimodal data is indistinguishable from N(0, 1), which is exactly what the ᾱt equation promised.
- The noise predictor - chapter 08's network, in miniature
- Two inputs (the noisy value, and the step t scaled to [0, 1]), one hidden layer of tanh units, one output: the predicted noise
H = 64 W1 = rng.standard_normal((H, 2)) * 0.5 # chapter 09: random init, small values b1 = np.zeros(H) W2 = rng.standard_normal(H) * 0.5 b2 = 0.0 def predict(x, t_frac): a1 = np.tanh(W1 @ np.vstack([x, t_frac]) + b1[:, None]) # hidden layer return W2 @ a1 + b2, a1
εθ(xt, t) as sixty-four hidden units. Feeding t in as an input is what lets one network learn to denoise at every noise level at once.
- Training - manufacture (noisy input, true noise) pairs and do mini-batch gradient descent on the squared error
- The gradient computation is chapter 09's backpropagation written out by hand for a two-layer net
lr = 1e-2 for step in range(4000): # mini-batch descent, chapter 17 x0 = sample_data(256) t = rng.integers(0, T, size=256) # a random step per example eps = rng.standard_normal(256) xt = np.sqrt(alpha_bar[t]) * x0 + np.sqrt(1 - alpha_bar[t]) * eps eps_hat, a1 = predict(xt, t / T) d = eps_hat - eps # the error to send backwards gW2 = a1 @ d / len(d) # backprop, chapter 09 in miniature: gb2 = d.mean() # output-layer gradients... da1 = np.outer(W2, d) * (1 - a1 ** 2) # ...delta for the hidden layer... gW1 = da1 @ np.vstack([xt, t / T]).T / len(d) gb1 = da1.mean(axis=1) # ...and its gradients W1 -= lr * gW1; b1 -= lr * gb1 # simultaneous update, chapter 02 W2 -= lr * gW2; b2 -= lr * gb2 # training loss falls from 4.1 to ~0.35 over the 4000 steps (under a second)
The whole training loop. Note what it never does: it never sees a "generated sample is good/bad" signal. It only ever learns to predict added noise, and generation quality follows from that alone.
- Generation - start from pure noise, run the reverse update 50 times
n = 4000 x = rng.standard_normal(n) # x_T: pure noise, no data in sight for t in range(T - 1, -1, -1): eps_hat, _ = predict(x, np.full(n, t / T)) x = (x - beta[t] / np.sqrt(1 - alpha_bar[t]) * eps_hat) / np.sqrt(alpha[t]) if t > 0: x = x + np.sqrt(beta[t]) * rng.standard_normal(n) # the fresh z (x < 0).mean() # 0.517 - half in each mode, as in the data x[x < 0].mean(), x[x >= 0].mean() # (-1.89, 1.87) - mode centres (true: -2, +2) x[x < 0].std(), x[x >= 0].std() # (0.43, 0.47) - mode widths (true: 0.3) (np.abs(x) < 1).mean() # 0.054 - the gap is nearly empty
Four thousand samples of standard normal noise, pushed backwards through the learned denoiser, come out bimodal: two clean modes at ±1.9 with an almost-empty gap between them. Slightly blurrier than the truth (widths 0.45 against 0.3) — about right for a sixty-four-unit network — but this is real generation: a distribution nothing in the sampling loop ever saw, reconstructed from noise.
- Scaling this up to images changes nothing conceptual
- x becomes a million-pixel array instead of one number; the two-layer net becomes a large U-Net or transformer; T becomes ~1000
- The schedule, the closed-form jump, the squared-error objective and the sampling loop are exactly the ones above
Conditioning and guidance
- Everything so far generates unconditionally - real systems generate from a prompt
- Conditioning - feed a text prompt (encoded by a language model) into the denoiser as an extra input, so εθ(xt, t, prompt) is steered towards matching images. The prompt attends into the denoiser via cross-attention - chapter 21
- Classifier-free guidance - run the denoiser with and without the prompt and extrapolate away from the unconditioned prediction. Exaggerates prompt adherence, at some cost in diversity
- Latent diffusion - run the whole process in a learned compressed space rather than on pixels, which is what made it cheap enough to use widely. Chapter 14's dimensionality-reduction argument, doing real work: denoise where the data actually varies
What they're used for
- Image and video generation and editing - inpainting, upscaling, style transfer
- Increasingly structures rather than pictures - protein backbone design (RFdiffusion and relatives), small molecules, materials
- AlphaFold3 replaced its final coordinate module with a diffusion process - chapter 24
- Anything where you want to sample from a complicated distribution rather than predict a single answer
- Regression (chapter 02) gives you the conditional mean; diffusion gives you draws from the whole conditional distribution
Summary
- Forward: destroy data with T small fixed Gaussian steps; a closed form jumps to any noise level directly
- Reverse: train a network to predict the added noise - squared error against a label you manufactured, so supervision is free
- Generate: start from pure noise and apply the reverse update T times, adding a little fresh noise each step
- The NumPy model above does all of this in ~40 lines - chapter 15's Gaussians, chapter 02's cost, chapter 08's network, chapter 09's backprop, chapter 17's mini-batches - and turns noise into a bimodal distribution it was never shown
- Conditioning brings in the prompt through cross-attention (chapter 21); latent diffusion moves the whole game into a compressed space (chapter 14)
- The deep idea: don't learn the hard leap from noise to data - learn one easy step, and take it many times