Inspect who experiences the errors
Use subgroup metrics carefully and avoid conclusions from tiny samples.
- Compare error types across slices
- Report denominators
- Recognize competing fairness objectives
An average can hide unequal behavior
A model may have similar overall accuracy while missing many more positive cases in one subgroup. Compare metrics relevant to the application, such as false-negative rates, false-positive rates, and calibration, rather than selecting only the most flattering summary.
Choose slices for a reason grounded in how the system is used. Data collection and labels may already contain inequities. Model evaluation cannot repair those assumptions by itself, and adding a sensitive attribute requires a legitimate purpose and careful handling.
Numbers need context
Small subgroup counts produce unstable estimates. Report the numerator and denominator and investigate representative failures. Do not rank groups or claim a systematic disparity from a handful of toy observations.
Fairness definitions can conflict when base rates differ. There is no single universal metric that settles every application. Document the intended use, affected people, acceptable trade-offs, and a review process.
A small experiment you can run.
Groups a and b are synthetic placeholders. Two positive examples per group are far too few for a reliable real-world conclusion.
rows = [("a", 1, 1), ("a", 1, 0), ("a", 0, 0),
("b", 1, 1), ("b", 1, 1), ("b", 0, 1)]
for group in ["a", "b"]:
positives = [(y, p) for g, y, p in rows if g == group and y == 1]
missed = sum(p == 0 for _, p in positives)
print(group, {"missed_positives": missed, "actual_positives": len(positives),
"false_negative_rate": missed/len(positives) if positives else None})
Save the file, open your terminal in that folder, and run python fairness-and-error-slices.py. Use python3 or py if required by your installation. Setup guide
The toy false-negative rates are 0.5 for a and 0 for b, based on two positives each.
Add denominators to a report.
- Calculate each group’s false-positive count and negative count.
- Compare with its false-negative result.
- Write one sentence explaining the sample-size limitation.
Compare with a suggested solution
Group a has no false positives among one negative; group b has one among one negative. These rates are extremely unstable. Report the counts and request more representative evidence before making a deployment claim.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
NIST AI Risk Management Framework