Learn AI/Deep learning
LESSON 18 / 36Intermediate 40 min with practice

Backpropagation without the mystery

Follow the chain rule through one neuron and check the derivative numerically.

WHAT YOU WILL LEARN
  • Apply the chain rule
  • Compare analytic and numerical gradients
  • Update parameters using old values

Work backward from the loss

Backpropagation is an efficient application of the chain rule. For a sigmoid neuron with squared error, the derivative of loss with respect to its weight combines the loss derivative, the sigmoid derivative, and the input.

Compute the forward pass first and keep the intermediate values. During the backward pass, calculate all gradients from those same parameter values before updating any of them. Updating halfway through can accidentally mix two different versions of the network.

Check a tiny case first

A finite-difference check estimates a derivative by comparing the loss at weight plus epsilon and weight minus epsilon. Dividing the difference by twice epsilon approximates the slope.

Numerical gradients are expensive for large models and depend on the choice of epsilon. They are excellent for checking a small implementation, especially before adding batches, multiple layers, or performance optimizations.

PUT THE IDEA INTO CODE

A small experiment you can run.

The check isolates one weight with the bias held fixed. Agreement is evidence that this derivative is implemented correctly, not that an entire network is correct.

gradients-and-backpropagation.py
import math
x, target, weight, bias = 0.7, 1., 0.3, -0.2
def prediction(w):
    return 1/(1+math.exp(-(w*x+bias)))
def loss(w):
    return (prediction(w)-target)**2
p = prediction(weight)
analytic = 2*(p-target)*p*(1-p)*x
epsilon = 1e-5
numeric = (loss(weight+epsilon)-loss(weight-epsilon))/(2*epsilon)
print("Analytic:", analytic, "Numeric:", numeric)
assert abs(analytic-numeric) < 1e-6
Copy code

Save the file, open your terminal in that folder, and run python gradients-and-backpropagation.py. Use python3 or py if required by your installation. Setup guide

What to expect

The original analytic and numerical gradients agree within 0.000001.

YOUR TURN

Check the bias derivative.

  1. Write a prediction function that varies the bias.
  2. Remove the final x factor from the analytic derivative.
  3. Compare with a central finite difference.
Compare with a suggested solution

The derivative of weight*x+bias with respect to bias is 1, while its derivative with respect to weight is x. All other chain-rule terms stay the same.

CHECK YOUR UNDERSTANDING

One idea to take with you.

When should parameters be updated?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

PyTorch: automatic differentiation