Precision, recall, and decision thresholds
See the trade-off between catching positives and creating false alerts.
- Build a confusion matrix
- Calculate precision and recall
- Select a threshold for a use case
Count the different mistakes
A true positive is a correctly flagged positive. A false positive is an ordinary case incorrectly flagged. A false negative is a positive the system missed. Precision is true positives divided by all predicted positives; recall is true positives divided by all actual positives.
Accuracy can conceal failure when one class dominates. A system that always predicts “ordinary” on a dataset with 99% ordinary cases has 99% accuracy but zero recall for the cases that may matter most.
A score is not a decision
A classifier often produces a score before making a binary decision. Raising the threshold usually reduces the number of positive predictions; lowering it usually increases them. The resulting precision and recall depend on the observed score distribution.
Choose the threshold on validation data using error costs, review capacity, and minimum requirements. When a denominator is zero, represent the metric as unavailable or follow an explicit convention. Do not silently turn an undefined result into a perfect score.
Move the threshold. See the trade-off.
Six invented predictions, one decision boundary. A higher score is classified as positive when it meets the threshold.
A small experiment you can run.
Use the interactive threshold lab to move the decision boundary on this same invented dataset. None of these scores is a live model benchmark.
actual = [1, 0, 1, 0, 1, 0]
scores = [0.95, 0.8, 0.65, 0.4, 0.3, 0.1]
for threshold in [0.3, 0.5, 0.9]:
predicted = [int(score >= threshold) for score in scores]
tp = sum(a == 1 and p == 1 for a, p in zip(actual, predicted))
fp = sum(a == 0 and p == 1 for a, p in zip(actual, predicted))
fn = sum(a == 1 and p == 0 for a, p in zip(actual, predicted))
precision = tp/(tp+fp) if tp+fp else None
recall = tp/(tp+fn) if tp+fn else None
print(threshold, {"precision": precision, "recall": recall})
Save the file, open your terminal in that folder, and run python precision-recall-and-thresholds.py. Use python3 or py if required by your installation. Setup guide
At threshold 0.5, precision and recall are both 2/3.
Choose a review threshold.
- Assume reviewers can inspect at most three flagged examples.
- Calculate how many cases each threshold flags.
- Choose a feasible threshold and report what it misses.
Compare with a suggested solution
At 0.5, three cases are flagged and two are correct; one actual positive is missed. At 0.9, only one case is flagged and two positives are missed. At 0.3, five flags exceed capacity. The choice depends on the stated constraints.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
scikit-learn: model evaluationscikit-learn: common pitfalls