Learn AI/Deep learning
LESSON 20 / 36Intermediate 30 min with practice

Overfitting, regularization, and early stopping

Choose a model checkpoint using validation behavior.

WHAT YOU WILL LEARN
  • Recognize diverging loss curves
  • Select a validation checkpoint
  • Keep final test data untouched

Memorizing is not generalizing

A flexible model can keep reducing training error while becoming worse on unseen examples. This is overfitting. It is especially easy with few examples, noisy labels, or repeated tuning against a familiar evaluation set.

Regularization limits or discourages overly complex solutions. Examples include weight penalties, smaller models, data augmentation appropriate to the task, and early stopping. Each changes the trade-off and must be validated.

Save the right checkpoint

Early stopping monitors a validation metric and retains the best observed parameter state. A patience setting allows a few non-improving steps before stopping. Save a copy of the weights, not a mutable reference that later updates will change.

The selected validation result is influenced by model choice and is therefore not an unbiased final estimate. Use a held-out test set after decisions are complete, and report uncertainty when the sample is small.

PUT THE IDEA INTO CODE

A small experiment you can run.

These are constructed learning curves, not training output from a real model. They isolate the checkpoint-selection decision.

overfitting-and-early-stopping.py
training = [0.8, 0.6, 0.4, 0.3, 0.2, 0.1]
validation = [0.9, 0.7, 0.5, 0.48, 0.53, 0.6]
best_epoch = min(range(len(validation)), key=validation.__getitem__)
print("Best checkpoint:", best_epoch + 1)
print("Training loss:", training[best_epoch])
print("Validation loss:", validation[best_epoch])
Copy code

Save the file, open your terminal in that folder, and run python overfitting-and-early-stopping.py. Use python3 or py if required by your installation. Setup guide

What to expect

The selected checkpoint is 4, with training loss 0.3 and validation loss 0.48.

YOUR TURN

Implement patience of two.

  1. Read validation losses in order.
  2. Reset a counter on improvement and increment it otherwise.
  3. Stop after two consecutive non-improvements while retaining the best epoch.
Compare with a suggested solution

The best checkpoint is epoch 4. Epochs 5 and 6 fail to improve, so training stops after epoch 6 with epoch 4 retained. The most recent weights are not necessarily the best weights.

CHECK YOUR UNDERSTANDING

One idea to take with you.

Which checkpoint should be kept here?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

PyTorch: automatic differentiation