22: Large Language Models
A note on this chapter
- Like chapters 20 and 21, this is an addition - there is no 2011 lecture behind it
- Chapter 20 gives the high-level tour of large language models; this chapter is the mathematics
- It assumes chapter 21 (attention) and leans constantly on chapters 06 and 09
- The punchline, stated up front
- A language model is chapter 06's classifier with a vocabulary-sized output, applied over and over
- Almost every equation below is one we have already seen, wearing bigger numbers
The language modelling problem
- A language model assigns a probability to a sequence of tokens
- Good sentences should get more probability than word salad
- That single requirement, pushed hard enough, is where everything else comes from
- A joint distribution over sequences is unmanageable directly
- Even a 10-token sentence over a 50,000-word vocabulary has 50,00010 possible values
- So factorize it with the chain rule of probability - no approximation involved
The chain rule: the probability of a sequence is the product of one next-token prediction per position. It turns "model language" into "predict the next token, repeatedly" — a supervised learning problem with the labels built in, which is the self-supervision idea from chapter 20.
- Each factor is a classification problem
- Input: the tokens so far. Output: a distribution over the vocabulary
- Exactly the multiclass setup of chapters 06 and 08, with |V| classes instead of 4
Tokens and embeddings
- Tokens
- Text is split into word-pieces from a fixed vocabulary V, typically 50,000-200,000 entries
- Common words are one token; rare words split into several ("tokenization" → token + ization)
- Fixes the output dimension, and nothing is ever out of vocabulary
- Embeddings
- Token i becomes a one-hot vector ei - all zeros with a 1 in position i, exactly the one-hot class vectors from chapter 08's multiclass section
- Multiply by a learned embedding matrix E, which is [|V| x d]
A one-hot vector times a matrix just selects row i of E — chapter 03's matrix-vector multiplication doing dictionary lookup. Each token gets a learned d-dimensional vector (d is typically 1,000-10,000), and tokens that behave similarly end up with similar vectors, because that is what reduces prediction error.
- Position is added to each embedding before any attention happens - chapter 21's positional encoding, needed for exactly the reason given there
- The sequence of embedded, position-tagged vectors then runs through the stack of transformer blocks from chapter 21
Autoregressive models
- The GPT family, and every current chat assistant, are autoregressive (AR) models
- They model the chain-rule factors directly: always predict the next token from what came before
- The output head
- The transformer turns the context into a vector; one final matrix maps it to |V| numbers - the logits z
- Softmax turns logits into a distribution
Softmax: exponentiate every logit and normalize. Positive, sums to 1, and the biggest logit gets the biggest share — the same function that turned scores into weights inside attention (chapter 21).
- Softmax is the sigmoid, grown up
- Chapter 06's sigmoid is exactly softmax over two classes with the second logit fixed at 0
Divide top and bottom by ez and chapter 06's hypothesis falls out. Binary classification was the |V| = 2 case all along.
import numpy as np
def softmax(z):
z = z - z.max() # subtract the max first - numerical safety,
e = np.exp(z) # and it changes nothing (top and bottom scale alike)
return e / e.sum()
softmax(np.array([1.3, 0.0]))[0] # 0.785835
1 / (1 + np.exp(-1.3)) # 0.785835 - the sigmoid, identically
The identity checked numerically: two-class softmax with the second logit at 0 is chapter 06's sigmoid.
- The training objective
- Maximise the log-probability of the training text; equivalently minimise the average negative log-likelihood
Cross-entropy loss. For two classes this is literally chapter 06's cost function — y log h + (1 − y) log(1 − h) is the |V| = 2 case — and it is maximum likelihood for the same reason: there the labels were modelled as Bernoulli, here as categorical. One objective, from spam filters to chat assistants.
vocab = ["mat", "dog", "moon", "sofa", "banana"] logits = np.array([3.2, 1.1, 0.4, 2.4, -1.0]) # "the cat sat on the ___" p = softmax(logits) # mat 0.607, sofa 0.273, dog 0.074, moon 0.037, banana 0.009 -np.log(p[0]) # 0.499 - the loss if the actual next word is "mat" -np.log(p[4]) # 4.699 - the loss if it is "banana": rare surprises cost a lot
One next-token prediction, scored. The loss is small when the model put probability on what actually happened, and large when it was surprised — averaging this over trillions of tokens is the whole of pretraining.
- Why one pass trains every position
- With chapter 21's causal mask, position t can only see positions 1 to t
- So a single forward pass over a T-token document yields T separate next-token predictions, each scored against the token that actually follows
- Every token of the corpus is a training example - no labelling budget, which is what chapters 10 and 11 spent so much care rationing
Masked language models
- The BERT family - and chapter 25's protein models - train differently
- Hide a random subset of tokens (typically 15%); predict each hidden token from everything else
M is the set of masked positions. Same softmax, same cross-entropy — the only changes are which positions are scored and what the model is allowed to look at.
- The crucial architectural difference is the attention mask
- No causal mask: every position attends to every position, in both directions
- "The ___ sat on the mat" - the right context is often what settles the answer
- What you give up
- The chain rule decomposition is gone - the products of masked conditionals do not multiply into a coherent p(sequence)
- So an MLM cannot generate text by construction; it fills in blanks
- And only the masked 15% of positions produce a training signal per pass, against 100% for the AR objective
- What you get
- Representations built from both directions at once - each position's vector summarises its full context
- Which is exactly what you want when the goal is embeddings for a downstream task rather than generation
Comparing the two objectives
| Autoregressive (GPT-style) | Masked (BERT-style) | |
|---|---|---|
| Predicts | the next token, from the left context | masked tokens, from both sides |
| Attention mask | causal (chapter 21) | none - fully bidirectional |
| Models p(sequence)? | yes, exactly, via the chain rule | no - conditionals don't assemble into a joint |
| Training signal per pass | every position | the masked ~15% |
| Can generate? | natively, one token at a time | not directly |
| Strongest at | generation: writing, code, dialogue | representations: embeddings, classification, retrieval |
| Examples | GPT, Claude, Gemini, Llama | BERT, RoBERTa, ESM (chapter 25) |
- Why the assistants are all autoregressive
- Assistants must generate, and AR models are exact generative models of the sequence distribution
- The denser training signal also pays at scale
- Why MLMs did not disappear
- When the product is an embedding - search, similarity, features for a small downstream model (chapter 10's workflow) - bidirectional context wins
- Protein models stayed masked for exactly this reason, plus one more: a protein is not written left to right, so there is no natural generation order to exploit - chapter 25
Generation, temperature and sampling
- Generation from an AR model is the chain rule run forwards
- Predict a distribution, pick a token, append it, repeat
- Each chosen token becomes context for the next prediction
- How to pick from the distribution
- Greedy - always take the argmax. Deterministic, and often repetitive and dull
- Sampling - draw from p. Faithful to the model, but its rare-token tail produces occasional nonsense
- Temperature - reshape the distribution before sampling
Divide the logits by a temperature τ before the softmax. τ → 0 approaches greedy; τ = 1 is the model's own distribution; τ > 1 flattens it towards uniform.
for tau in (0.5, 1.0, 2.0):
pt = softmax(logits / tau)
# tau=0.5: top prob 0.819 - sharpened, nearly greedy
# tau=1.0: top prob 0.607 - the model as trained
# tau=2.0: top prob 0.419 - flattened, more adventurous
The same five logits from above at three temperatures. One knob trades reliability against variety, with no retraining.
- In practice the tail is also cut before sampling (top-k or top-p), which removes most of the nonsense at little cost
Measuring a language model: perplexity
The exponential of the average loss. Interpretation: the model is, on average, as uncertain as if it were choosing uniformly between PPL tokens. A model that knows nothing about a 10-token vocabulary scores exactly 10; a good English model scores under 10 on a 100,000-token vocabulary.
- This is the single-real-number evaluation metric of chapter 11, for language models
- Lower is better; 1 would be a model that is never surprised
- The chapter 10 warning applies with full force
- Perplexity and benchmark scores are only meaningful on text the model has not trained on
- When the training set is a crawl of the internet, guaranteeing that is genuinely hard - contamination is the field's version of evaluating on the training set
Training at scale
- The optimization is the one we already know
- Mini-batch gradient descent (chapter 17) on the cross-entropy loss, over trillions of tokens
- In practice Adam (chapter 20) plus weight decay, which is chapter 07's regularization under its other name
- Backpropagation through the transformer stack is chapter 09's algorithm; the residual connections exist to keep its gradients alive (chapter 21)
- Scaling laws
- The loss falls as a power law in parameters N, data D and compute
An empirical regularity, not a theorem: straight lines on log-log plots over many orders of magnitude, with matching laws for data and compute. Predictable enough to budget a training run in advance — and the same fitting-a-curve-to-observations exercise as chapter 04's polynomial regression, applied to the models themselves.
- The practical corollary
- For a fixed compute budget there is an optimal balance of model size and data - training a smaller model on more tokens often beats a bigger model trained short
- Getting this trade-off right is a bias/variance argument (chapter 10) conducted with power laws
- From raw model to assistant
- Everything above produces a next-token predictor, not a helpful conversationalist
- The post-training stages - supervised fine-tuning, preference learning, reasoning training - are covered in chapter 20 and not repeated here
Summary
- Factorize p(sequence) with the chain rule; each factor is multiclass classification over the vocabulary
- Autoregressive models predict the next token under a causal mask - exact generative models, every position a training example
- Masked models predict hidden tokens from both directions - better representations, no generation
- Both use the same loss: softmax plus cross-entropy, which is chapter 06's maximum-likelihood cost with |V| classes; the sigmoid is its two-class special case
- Temperature reshapes the output distribution at generation time; perplexity is the exponential of the loss, and chapter 10's held-out discipline decides whether it means anything
- Training is chapter 17's mini-batch descent with chapter 09's backpropagation, at a scale set by empirical power laws
- The through-line: nothing in this chapter required a new idea beyond attention - the 2011 toolkit, scaled up, is the modern language model