holehouse.org Blog Machine learning notes

10: Advice for applying Machine Learning

Deciding what to try next

Debugging a learning algorithm

J(θ)=12m[∑i=1m(hθ(x(i))−y(i))2+λ∑j=1nθj2] The original slide sums the penalty to m; it runs to n, the number of features — m is already used by the sum over training examples.

Evaluating a hypothesis

A typical 70:30 split — the first 70% becomes the training set, the rest the test set
SizePriceSet
2104400training
1600330training
2400369training
1416232training
3000540training
1985300training
1534315training
1427199test
1380212test
1494243test

The training set gives (x(1), y(1)) … (x(m), y(m)); the test set gives (xtest(1), ytest(1)) … (xtest(mtest), ytest(mtest)).

Model selection and training validation test sets

Diagnosis - bias vs. variance

Regularization and bias/variance

Model:hθ(x)=θ0+θ1x+θ2x2+θ3x3+θ4x4 J(θ)=12m∑i=1m(hθ(x(i))−y(i))2+λ2m∑j=1nθj2 As above, the penalty sums to n rather than m on the original slide.

Learning curves

What to do next (revisited)