holehouse.org Blog Machine learning notes

25: ESM2 and ESM-C

A note on this chapter

Proteins as language

The objective: masked residues

J(θ)=−1|M|∑i∈Mlogpθ(xi∣x∖M) Chapter 22's masked loss, verbatim — only the vocabulary changed: ~33 tokens (20 amino acids plus specials) instead of 100,000 word-pieces. M is the masked ~15% of positions; following BERT, most are replaced by a mask token, a few by random residues, a few left alone.

ESM2

ESM-C

Comparing ESM2 and ESM-C

Same objective, same interface — the differences are data, recipe and efficiency.
ESM2 (2022-23)ESM-C (2024)
DeveloperMeta AIEvolutionaryScale
Objectivemasked residues (chapter 22 MLM)identical
Sizes8M – 15B parameters300M, 600M, 6B
Training dataUniRef, tens of millions of sequencesseveral-fold larger, heavy metagenomic fraction
Architecturetransformer encoder, rotary positionssame skeleton, modernised recipe
Rule of thumbthe established baselineESM2-quality embeddings at ~½ to ⅕ the size
Structure headESMFold (on the 3B model)none - embeddings are the product

Using the models: the maths

z=1r∑i=1rhi Mean-pooling the per-residue vectors hi. The fixed-length z then feeds a small supervised model — chapter 06's logistic regression on a few hundred labelled examples is often enough, which changes what a small lab can attempt. This is the "learned features replacing hand-built features" thread that runs from chapter 08 through chapter 16, completed.
s(mut)=logpθ(xi=mut∣x∖i)−logpθ(xi=wt∣x∖i) The masked-marginal score: how much less plausible does the model find the mutant than what evolution kept? Strongly negative predicts damage. No labelled variant data is used at any point — it is chapter 15's anomaly-detection logic (flag what the density model finds improbable) with a learned p.
import numpy as np

# toy numbers: the model's masked distribution at one buried position,
# where the wildtype is leucine (L)
p = {"L": 0.62, "I": 0.21, "V": 0.11, "P": 0.002}

np.log(p["I"] / p["L"])   # -1.08 - conservative substitution: mildly suspect
np.log(p["V"] / p["L"])   # -1.73 - similar, a little worse
np.log(p["P"] / p["L"])   # -5.74 - proline in a buried helix: strongly deleterious

The scoring rule on illustrative numbers: chemically similar residues score near zero, and the substitution that breaks the local structure scores far below. Real use sums this over every mutated position.

pPPL=exp(−1r∑i=1rlogpθ(xi∣x∖i)) Between 1 (certain) and 20 (uniform over the amino acids). Well-modelled protein families sit far below 20, and a sequence's pseudo-perplexity tracks how natural the model finds it — low pPPL designs are likelier to fold.

Caveats

Summary