24: AlphaFold2
A note on this chapter
- An addition, like chapters 20-23 - and the one furthest from the 2011 course in subject, but not in machinery
- Attention (chapter 21) appears here in four distinct costumes; the masked-token objective (chapter 22) turns up as an auxiliary loss; diffusion (chapter 23) is where the successor went
- This is a description of how AlphaFold2 works, piece by piece - not of how to run it
The problem
- A protein is a chain of amino acids that folds into a specific three-dimensional structure
- The structure largely determines the function, and the sequence largely determines the structure
- Predicting structure from sequence was an open problem for fifty years - experimental determination is slow and expensive
- What AlphaFold2 (DeepMind, 2020) did
- For most single-domain proteins, predictions competitive with experiment
- Assessed blind at CASP14 - the community's held-out test set, structures solved but not yet published
- Chapter 10 discipline at the scale of a whole field: the assessment being genuinely blind is why the result was believed immediately rather than argued about for years
- Framed in this course's terms
- Input x: a sequence (length r, alphabet of 20). Output y: 3D coordinates for every atom
- A supervised learning problem (chapter 01) - with ~170,000 known structures as the labelled training set
The key input: MSAs and coevolution
- The single most important input is not the sequence alone but a multiple sequence alignment (MSA)
- The same protein collected from many organisms and aligned column-by-column - often thousands of variants
- Why that carries structural information
- If two positions are in contact in the folded structure, a mutation at one tends to be compensated by a mutation at the other - or the protein breaks and the organism is not around to be sequenced
- So contacting positions co-vary down the columns of the alignment
- Statistical dependence between columns (the correlations chapter 14's covariance matrix measures) is a fingerprint of physical proximity - evolution has been running the mutagenesis experiment for us
- Pre-AlphaFold methods computed these correlations explicitly and predicted contact maps from them
- AlphaFold2 instead hands the raw alignment to the network and lets it learn what to extract - the feature-learning argument of chapters 08 and 16, applied to evolution's data
The architecture at a glance
- Two internal representations, refined together, then turned into coordinates
- MSA representation - a [s × r] array: one row per aligned sequence, one column per residue
- Pair representation - a [r × r] array: one cell per pair of residues, holding what the model currently believes about their relationship (distance, orientation)
- The design principle
- Keep a "what does each sequence say about each position" view and a "what is the relationship between every pair of positions" view, and let them inform each other repeatedly
- Structure lives in the pair view; evidence lives in the MSA view
The Evoformer
- 48 identical blocks, each refining both representations - four distinct mechanisms per block
- 1) Row-wise attention on the MSA, biased by the pair representation
- Within one aligned sequence, each residue attends to the others - chapter 21's self-attention along the row
- With one addition: a learned bias from the pair representation is added to the attention scores
Chapter 21's scaled dot-product score with one extra term: bij comes from the pair representation's cell (i, j). If the model currently believes residues i and j are close, their attention is boosted — the pair view steers where the MSA view looks.
- 2) Column-wise attention on the MSA
- Down a column - the same position across organisms - letting evidence flow between sequences
- Rows and columns alternating is how information gets everywhere without attending over the full s × r array at once (chapter 21's O(n²) cost, managed)
- 3) Outer product mean: MSA → pair
- For each pair of columns (i, j), take the outer product of their representations averaged over sequences - a learned generalisation of "measure the covariation between column i and column j"
- The coevolution statistic that older methods hand-crafted, rediscovered as a differentiable layer
- 4) Triangle updates on the pair representation
- Distances obey constraints: if i is close to k and k is close to j, then the (i, j) distance is bounded - the triangle inequality
- So the pair cell (i, j) is updated by attending over every third residue k, along the triangle's edges (i, k) and (j, k) - both a multiplicative update and a triangle-attention variant
- Geometry is never imposed; the update pattern just makes geometric consistency easy to learn
- Plus the transformer usuals from chapter 21 - residual connections, normalisation, and feed-forward transitions after each attention
The structure module
- The stage that turns representations into an actual fold - 8 blocks, sharing weights
- Each residue is given a rigid frame: a rotation and a translation - its backbone treated as a small rigid body floating in space
- All frames start at the origin, and the module iteratively moves them
- Invariant Point Attention (IPA) - the third costume attention wears here
- Attention scores get an extra geometric term: each residue emits query and key points in its own frame, and the score depends on the distance between them in global space
- Built so the result is unchanged if the whole structure is rotated or translated - the answer should not depend on where the molecule happens to sit, so the invariance is baked into the architecture rather than learned
- The same reasoning as feature scaling in chapter 04: don't make the network spend capacity learning something you can guarantee by construction
- After the frames settle, predicted torsion angles place the side-chain atoms, giving all-atom coordinates
Training, losses and recycling
- The main loss: FAPE (frame-aligned point error)
- For every residue's frame, express every atom's predicted and true positions in that local frame, and penalise the distance between them - averaged over all frames and atoms
- Comparing in local frames makes the loss invariant to global rotation/translation, and clamping large errors keeps single disasters from dominating - the same robustness instinct as capping a cost function
- Auxiliary losses, trained jointly
- A distogram head: predict the distance distribution for every pair - multiclass classification (chapters 08/22) over distance bins, supervising the pair representation directly
- A masked-MSA head: hide MSA entries and predict them - literally chapter 22's masked-language objective, keeping the MSA representation honest
- The confidence heads (next section), trained to predict the model's own error
- Recycling
- The final representations and predicted structure are fed back in as inputs and the whole network runs again - typically 3 extra passes
- An iterative-refinement loop, the gradient-descent instinct of chapter 02 applied at the scale of whole predictions: start somewhere, improve repeatedly
- Self-distillation
- After training on the ~170k known structures, predict structures for ~350k unlabelled sequences, keep the confident ones, and retrain on the enlarged set
- Manufacturing labels from unlabelled data - the self-supervision theme of chapter 20 again, with the model's own confidence as the filter
Confidence: pLDDT and PAE
- The part that gets underrated - AlphaFold2 tells you when to trust it
- pLDDT - per-residue confidence, 0-100, predicted by a head trained to estimate the local accuracy of the model's own output
- PAE - predicted aligned error: for every pair of residues, the expected error in the position of one when aligned on the frame of the other. This is what you need for judging whether the relative placement of two domains is trustworthy
- Both are learned predictions of the model's own error, and they are well calibrated
- A calibrated confidence estimate is what turns a prediction into something usable in practice - chapter 06 made the same point about h(x) being a probability rather than a bare label
Limits
- It predicts a structure, and does that well for proteins that have one
- It is not a folding simulation; it says little about dynamics or the conformational ensemble
- Low pLDDT frequently means genuinely disordered rather than merely uncertain - a low-confidence region is itself biological signal
- Effects of point mutations, binding, and conditions are largely outside what it was trained to do
- Shallow MSAs hurt - the coevolution signal is the fuel, and orphan sequences don't carry it (which is where chapter 25's single-sequence models come in)
After AlphaFold2
- AlphaFold-Multimer extended the recipe to protein complexes
- AlphaFold3 (2024) restructured it
- The structure module is replaced by a diffusion model over raw atom coordinates - chapter 23's noise-to-structure loop, conditioned on the trunk's representations
- Coverage extends to nucleic acids, ligands and modifications; the price is sampling multiple draws and ranking by confidence
- That a diffusion head could replace the carefully hand-built structure module is the chapter 23 thesis in action: given good conditioning, "denoise your way to the answer" is a remarkably general decoder
Summary
- The problem: sequence in, all-atom 3D structure out - supervised learning against the protein data bank, judged blind at CASP
- The decisive input is the MSA: contacting residues co-vary across evolution, so alignment statistics encode geometry
- Two representations - MSA [s × r] and pair [r × r] - refined by 48 Evoformer blocks: row and column attention (pair-biased), outer-product mean, and triangle updates that make geometric consistency learnable
- The structure module poses each residue as a rigid frame and moves them with Invariant Point Attention - rotation/translation invariance by construction
- Trained with FAPE plus distogram, masked-MSA (chapter 22's objective) and confidence heads; recycled ~3 times; boosted by self-distillation
- pLDDT and PAE are calibrated self-error estimates - the feature that makes the predictions usable
- AlphaFold3 swapped the structure module for a diffusion decoder (chapter 23) - the architecture is modular enough that its final act could be replaced wholesale