Stanford Machine Learning
The following notes represent a complete, stand alone interpretation of Stanford's machine learning course presented by Professor Andrew Ng and originally posted on the ml-class.org website during the fall 2011 semester. The topics covered are shown below, although for a more detailed summary see lecture 19. The code examples throughout use Python with NumPy rather than the programming environment the original course taught.
All diagrams are my own or are directly taken from the lectures, full credit to Professor Ng for a truly exceptional lecture course.
What are these notes?
Originally written as a way for me personally to help solidify and document the concepts, these notes have grown into a reasonably complete block of reference material spanning the course in its entirety in just over 40,000 words and a lot of diagrams! The target audience was originally me, but more broadly, can be someone familiar with programming although no assumption regarding statistics, calculus or linear algebra is made. We go from the very introduction of machine learning to neural networks, recommender systems and even pipeline design. The one thing I will say is that a lot of the later topics build on those of earlier sections, so it's generally advisable to work through in chronological order.
The notes were written in Evernote and exported to HTML automatically, which left the original pages with a lot of markup that was never meant to be read. This version rewrites that markup by hand-built conversion: the words are unchanged, but the structure underneath them is now clean, semantic HTML.
How can you help!?
If you notice errors or typos, inconsistencies or things that are unclear please tell me and I'll update them. It would be hugely appreciated!
You can find me at alex[AT]holehouse[DOT]org
A changelog can be found here — it lists both the original corrections and everything that changed in this rewrite.
Contents
Each chapter is tagged with where it came from. 2011 course marks the original notes from the Stanford course: the text is unchanged apart from corrections, although the code examples were rewritten in Python and NumPy in 2026. 2026 addition marks chapters written for this edition, covering material that did not exist or was not part of the course in 2011.
- 01: Introduction 2011 course What machine learning is, and the difference between supervised and unsupervised learning.
- 02: Linear Regression with One Variable 2011 course Fitting a straight line: the hypothesis, the cost function, and gradient descent from first principles.
- 03: Linear Algebra — Review 2011 course Matrices, vectors, multiplication, inverse and transpose: the notation the rest of the course leans on.
- 04: Linear Regression with Multiple Variables 2011 course Many features at once: vectorized gradient descent, feature scaling, learning rates, polynomial features and the normal equation.
- 05: Probability and Bayes' Rule 2026 addition The probability the course assumes but never teaches: Bayes' rule, maximum likelihood, and naive Bayes. Replaces the programming chapter that was never written.
- 06: Logistic Regression 2011 course Classification with the sigmoid: decision boundaries, the log-loss cost, and one-vs-all for more than two classes.
- 07: Regularization 2011 course Overfitting, and how a penalty on the parameters tames it in both linear and logistic regression.
- 08: Neural Networks — Representation 2011 course How a network is wired: layers, activations, forward propagation, and why it can compute non-linear functions.
- 09: Neural Networks — Learning 2011 course Training a network: the cost function, backpropagation, gradient checking and random initialization.
- 10: Advice for applying machine learning techniques 2011 course Diagnosing what is wrong with a model: train, validation and test sets, bias versus variance, and learning curves.
- 11: Machine Learning System Design 2011 course Building a system end to end, using spam classification: what to prioritise, error metrics for skewed classes, and precision versus recall.
- 12: Support Vector Machines 2011 course Large-margin classification, the kernel trick for non-linear boundaries, and when to prefer an SVM over logistic regression.
- 13: Clustering 2011 course Unsupervised learning with K-means: the algorithm, its objective, and choosing the number of clusters.
- 14: Dimensionality Reduction 2011 course Principal component analysis for compressing and visualising data, and how many components to keep.
- 15: Anomaly Detection 2011 course Modelling normal data with a Gaussian and flagging what does not fit, including the multivariate case.
- 16: Recommender Systems 2011 course Content-based and collaborative filtering, and low-rank matrix factorization for predicting ratings.
- 17: Large Scale Machine Learning 2011 course Learning from very large datasets: stochastic and mini-batch gradient descent, online learning, and map-reduce.
- 18: Application Example — Photo OCR 2011 course A complete pipeline: sliding windows, artificial data synthesis, and ceiling analysis to decide what to improve next.
- 19: Course Summary 2011 course A short recap of everything the original course covered.
- 20: Modern Deep Learning 2026 addition What changed after 2011: ReLU, initialization, Adam, dropout and batch normalization, and how today's assistants are trained.
- 21: Attention 2026 addition Attention as a soft lookup, self-attention, positional encoding, multi-head attention, and the transformer block.
- 22: Large Language Models 2026 addition Tokens and embeddings, autoregressive and masked objectives, sampling, perplexity, and training at scale.
- 23: Diffusion Models 2026 addition Destroying data with noise and learning to reverse it, with a complete diffusion model in NumPy.
- 24: AlphaFold2 2026 addition How protein structure prediction works: MSAs and coevolution, the Evoformer, the structure module, and confidence measures.
- 25: ESM2 and ESM-C 2026 addition Protein language models: masked residue prediction, what the embeddings capture, and how to use them in practice.
- 26: Convolutional Neural Networks 2026 addition Convolution, pooling and weight sharing for image data, the ImageNet moment, and how convnets relate to transformers.
- Appendix 1: Python and NumPy 2026 addition The NumPy primer behind the code examples: arrays, shapes, broadcasting and vectorization.