holehouse.org Blog Machine learning notes

22: Large Language Models

A note on this chapter

The language modelling problem

p(w1,w2,…,wT)=∏t=1Tp(wt∣w1,…,wt−1) The chain rule: the probability of a sequence is the product of one next-token prediction per position. It turns "model language" into "predict the next token, repeatedly" — a supervised learning problem with the labels built in, which is the self-supervision idea from chapter 20.

Tokens and embeddings

xi=ETei A one-hot vector times a matrix just selects row i of E — chapter 03's matrix-vector multiplication doing dictionary lookup. Each token gets a learned d-dimensional vector (d is typically 1,000-10,000), and tokens that behave similarly end up with similar vectors, because that is what reduces prediction error.

Autoregressive models

p(wt=v∣w<t)=ezv∑v′=1|V|ezv′ Softmax: exponentiate every logit and normalize. Positive, sums to 1, and the biggest logit gets the biggest share — the same function that turned scores into weights inside attention (chapter 21).
ezez+e0=11+e−z=g(z) Divide top and bottom by ez and chapter 06's hypothesis falls out. Binary classification was the |V| = 2 case all along.
import numpy as np

def softmax(z):
    z = z - z.max()            # subtract the max first - numerical safety,
    e = np.exp(z)              # and it changes nothing (top and bottom scale alike)
    return e / e.sum()

softmax(np.array([1.3, 0.0]))[0]   # 0.785835
1 / (1 + np.exp(-1.3))             # 0.785835 - the sigmoid, identically

The identity checked numerically: two-class softmax with the second logit at 0 is chapter 06's sigmoid.

J(θ)=−1T∑t=1Tlogpθ(wt∣w<t) Cross-entropy loss. For two classes this is literally chapter 06's cost function — y log h + (1 − y) log(1 − h) is the |V| = 2 case — and it is maximum likelihood for the same reason: there the labels were modelled as Bernoulli, here as categorical. One objective, from spam filters to chat assistants.
vocab  = ["mat", "dog", "moon", "sofa", "banana"]
logits = np.array([3.2, 1.1, 0.4, 2.4, -1.0])   # "the cat sat on the ___"
p = softmax(logits)
# mat 0.607, sofa 0.273, dog 0.074, moon 0.037, banana 0.009

-np.log(p[0])   # 0.499 - the loss if the actual next word is "mat"
-np.log(p[4])   # 4.699 - the loss if it is "banana": rare surprises cost a lot

One next-token prediction, scored. The loss is small when the model put probability on what actually happened, and large when it was surprised — averaging this over trillions of tokens is the whole of pretraining.

Masked language models

J(θ)=−1|M|∑t∈Mlogpθ(wt∣w∖M) M is the set of masked positions. Same softmax, same cross-entropy — the only changes are which positions are scored and what the model is allowed to look at.

Comparing the two objectives

The same architecture and the same cross-entropy loss — the objective and the mask are the entire difference.
Autoregressive (GPT-style)Masked (BERT-style)
Predictsthe next token, from the left contextmasked tokens, from both sides
Attention maskcausal (chapter 21)none - fully bidirectional
Models p(sequence)?yes, exactly, via the chain ruleno - conditionals don't assemble into a joint
Training signal per passevery positionthe masked ~15%
Can generate?natively, one token at a timenot directly
Strongest atgeneration: writing, code, dialoguerepresentations: embeddings, classification, retrieval
ExamplesGPT, Claude, Gemini, LlamaBERT, RoBERTa, ESM (chapter 25)

Generation, temperature and sampling

pv∝ezv/τ Divide the logits by a temperature τ before the softmax. τ → 0 approaches greedy; τ = 1 is the model's own distribution; τ > 1 flattens it towards uniform.
for tau in (0.5, 1.0, 2.0):
    pt = softmax(logits / tau)
    # tau=0.5: top prob 0.819  - sharpened, nearly greedy
    # tau=1.0: top prob 0.607  - the model as trained
    # tau=2.0: top prob 0.419  - flattened, more adventurous

The same five logits from above at three temperatures. One knob trades reliability against variety, with no retraining.

Measuring a language model: perplexity

PPL=exp(−1T∑t=1Tlogpθ(wt∣w<t)) The exponential of the average loss. Interpretation: the model is, on average, as uncertain as if it were choosing uniformly between PPL tokens. A model that knows nothing about a 10-token vocabulary scores exactly 10; a good English model scores under 10 on a 100,000-token vocabulary.

Training at scale

L(N)≈(NcN)αN An empirical regularity, not a theorem: straight lines on log-log plots over many orders of magnitude, with matching laws for data and compute. Predictable enough to budget a training run in advance — and the same fitting-a-curve-to-observations exercise as chapter 04's polynomial regression, applied to the models themselves.

Summary