Loss and gradient descent
See how repeated small updates can improve a prediction.
- Define squared error
- Calculate a simple derivative
- Observe learning-rate behavior
Turn an error into an objective
For a prediction y_hat and target y, squared error is (y_hat - y) squared. Averaging it over examples gives mean squared error. Squaring penalizes large misses more strongly and makes positive and negative errors contribute positively.
For the simple model y_hat = weight times x, the derivative of squared error with respect to weight is 2 times x times (weight times x minus y). The derivative describes the local direction and rate of loss change, not the best possible final weight.
Update carefully
Gradient descent subtracts a learning rate times the derivative from each parameter. Too small a learning rate can make progress slow; too large can overshoot or diverge. A decreasing training loss does not prove good generalization.
Check a baseline, inspect the loss curve, and evaluate unseen examples. Feature scale affects the gradients, so a learning rate that works for inputs near 1 can fail when inputs are near 10,000.
A small experiment you can run.
The dataset follows exactly y = 2x, so a weight close to 2 is expected. Real data usually contains noise and may not follow a straight line.
xs, ys = [1., 2., 3.], [2., 4., 6.]
weight, rate = 0., 0.05
for step in range(60):
gradient = sum(2*x*(weight*x-y) for x, y in zip(xs, ys)) / len(xs)
weight -= rate * gradient
loss = sum((weight*x-y)**2 for x, y in zip(xs, ys)) / len(xs)
print("Weight:", round(weight, 4), "Loss:", round(loss, 8))
print("Prediction for 4:", round(weight * 4, 3))
Save the file, open your terminal in that folder, and run python loss-and-gradient-descent.py. Use python3 or py if required by your installation. Setup guide
The original weight approaches 2 and the prediction for 4 approaches 8.
Find an unstable learning rate.
- Run with rates 0.005, 0.05, and 0.3.
- Record the final weight and loss.
- Explain why a larger step can move away from the solution.
Compare with a suggested solution
With this feature scale, 0.05 converges rapidly, 0.005 moves more slowly, and 0.3 overshoots repeatedly. Compare loss rather than assuming more aggressive updates are better.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
Python math module