Learn AI/Machine learning
LESSON 15 / 36Intermediate 35 min with practice

Precision, recall, and decision thresholds

See the trade-off between catching positives and creating false alerts.

WHAT YOU WILL LEARN
  • Build a confusion matrix
  • Calculate precision and recall
  • Select a threshold for a use case

Count the different mistakes

A true positive is a correctly flagged positive. A false positive is an ordinary case incorrectly flagged. A false negative is a positive the system missed. Precision is true positives divided by all predicted positives; recall is true positives divided by all actual positives.

Accuracy can conceal failure when one class dominates. A system that always predicts “ordinary” on a dataset with 99% ordinary cases has 99% accuracy but zero recall for the cases that may matter most.

A score is not a decision

A classifier often produces a score before making a binary decision. Raising the threshold usually reduces the number of positive predictions; lowering it usually increases them. The resulting precision and recall depend on the observed score distribution.

Choose the threshold on validation data using error costs, review capacity, and minimum requirements. When a denominator is zero, represent the metric as unavailable or follow an explicit convention. Do not silently turn an undefined result into a perfect score.

INTERACTIVE LABRUNS IN YOUR BROWSER

Move the threshold. See the trade-off.

Six invented predictions, one decision boundary. A higher score is classified as positive when it meets the threshold.

0.95Actual +
0.80Actual −
0.65Actual +
0.40Actual −
0.30Actual +
0.10Actual −
Precision67%
Recall67%
False alerts1
Missed positives1
Synthetic teaching data. Undefined precision is shown as N/A.
PUT THE IDEA INTO CODE

A small experiment you can run.

Use the interactive threshold lab to move the decision boundary on this same invented dataset. None of these scores is a live model benchmark.

precision-recall-and-thresholds.py
actual = [1, 0, 1, 0, 1, 0]
scores = [0.95, 0.8, 0.65, 0.4, 0.3, 0.1]
for threshold in [0.3, 0.5, 0.9]:
    predicted = [int(score >= threshold) for score in scores]
    tp = sum(a == 1 and p == 1 for a, p in zip(actual, predicted))
    fp = sum(a == 0 and p == 1 for a, p in zip(actual, predicted))
    fn = sum(a == 1 and p == 0 for a, p in zip(actual, predicted))
    precision = tp/(tp+fp) if tp+fp else None
    recall = tp/(tp+fn) if tp+fn else None
    print(threshold, {"precision": precision, "recall": recall})
Copy code

Save the file, open your terminal in that folder, and run python precision-recall-and-thresholds.py. Use python3 or py if required by your installation. Setup guide

What to expect

At threshold 0.5, precision and recall are both 2/3.

YOUR TURN

Choose a review threshold.

  1. Assume reviewers can inspect at most three flagged examples.
  2. Calculate how many cases each threshold flags.
  3. Choose a feasible threshold and report what it misses.
Compare with a suggested solution

At 0.5, three cases are flagged and two are correct; one actual positive is missed. At 0.9, only one case is flagged and two positives are missed. At 0.3, five flags exceed capacity. The choice depends on the stated constraints.

CHECK YOUR UNDERSTANDING

One idea to take with you.

If no examples are predicted positive, precision is…

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

scikit-learn: model evaluationscikit-learn: common pitfalls