holehouse.org Blog Machine learning notes

06: Logistic Regression

Classification

Hypothesis representation

Interpreting hypothesis output

Decision boundary

Decision boundary

Non-linear decision boundaries

Cost function for logistic regression

Training set:{(x(1),y(1)),(x(2),y(2)),…,(x(m),y(m))} x∈[x0x1⋮xn]x0=1,y∈{0,1} hθ(x)=11+e−θTx m examples, each a column vector of n+1 features.
J(θ)=1m∑i=1m12(hθ(x(i))−y(i))2

A convex logistic regression cost function

Cost(hθ(x),y)=
−log(hθ(x))ify=1
−log(1−hθ(x))ify=0

Simplified cost function and gradient descent

J(θ)=1m∑i=1mCost(hθ(x(i)),y(i))
Cost(hθ(x),y)=
−log(hθ(x))ify=1
−log(1−hθ(x))ify=0
y is always 0 or 1.
J(θ)=−1m[∑i=1my(i)loghθ(x(i))+(1−y(i))log(1−hθ(x(i)))] Because y is always 0 or 1, one of the two terms is always zero — so this single expression covers both cases.
def J(theta, X, y):
    h = g(X @ theta)
    return -(y * np.log(h) + (1 - y) * np.log(1 - h)).mean()

X = np.array([[1, 1], [1, 2], [1, 3], [1, 4]], dtype=float)
y = np.array([0, 0, 1, 1], dtype=float)

J(np.zeros(2), X, y)   # 0.6931 = log(2): with theta = 0 every h is 0.5

The compressed cost, vectorized. With θ = 0 the hypothesis says 0.5 for everything, and the cost is exactly log 2 per example.

How to minimize the logistic regression cost function

Repeat { θj≔θj−α∑i=1m(hθ(x(i))−y(i))xj(i) } simultaneously update all θj
theta = np.zeros(2)
for _ in range(20000):
    h = g(X @ theta)
    theta = theta - 0.5 * X.T @ (h - y) / len(y)

theta                  # [-25.6, 10.3]
-theta[0] / theta[1]   # 2.48 -> the decision boundary sits between x = 2 and x = 3
np.round(g(X @ theta), 3)   # [0., 0.007, 0.995, 1.]  - the fit

The same update as linear regression, but h is the sigmoid. The boundary lands at x ≈ 2.5 — between the last y = 0 example and the first y = 1.

Advanced optimization

J(θ) ∂∂θjJ(θ) for j = 0, 1, …, n
Repeat {θj≔θj−α∂∂θjJ(θ)}

Using advanced cost minimization algorithms

θ=[θ1θ2] J(θ)=(θ1−5)2+(θ2−5)2 ∂∂θ1J(θ)=2(θ1−5) ∂∂θ2J(θ)=2(θ2−5)
from scipy.optimize import minimize

def cost_function(theta):
    jval = ((theta - 5) ** 2).sum()
    gradient = 2 * (theta - 5)
    return jval, gradient

res = minimize(cost_function, np.zeros(2), jac=True)
res.x   # [5., 5.]  - the minimum, exactly where the derivatives hit zero

The worked example above, run for real: minimize drives both parameters to 5 without us ever picking a learning rate.

def cost_function(theta):
    ...
    return jval, gradient
from scipy.optimize import minimize

initial_theta = np.zeros(2)                  # initialize the theta values
res = minimize(cost_function, initial_theta, # run the algorithm
               jac=True, options={'maxiter': 100})
opt_theta, function_val = res.x, res.fun
THETA=[θ0θ1⋮θn]
def cost_function(theta):

    jval = [code to compute J(θ)];

    gradient[0] = [code to compute ∂∂θ0J(θ)];

    gradient[1] = [code to compute ∂∂θ1J(θ)];
    ⋮
    gradient[n] = [code to compute ∂∂θnJ(θ)];

Multiclass classification problems