Learn AI/Responsible AI
LESSON 30 / 36Intermediate 35 min with practice

Inspect who experiences the errors

Use subgroup metrics carefully and avoid conclusions from tiny samples.

WHAT YOU WILL LEARN
  • Compare error types across slices
  • Report denominators
  • Recognize competing fairness objectives

An average can hide unequal behavior

A model may have similar overall accuracy while missing many more positive cases in one subgroup. Compare metrics relevant to the application, such as false-negative rates, false-positive rates, and calibration, rather than selecting only the most flattering summary.

Choose slices for a reason grounded in how the system is used. Data collection and labels may already contain inequities. Model evaluation cannot repair those assumptions by itself, and adding a sensitive attribute requires a legitimate purpose and careful handling.

Numbers need context

Small subgroup counts produce unstable estimates. Report the numerator and denominator and investigate representative failures. Do not rank groups or claim a systematic disparity from a handful of toy observations.

Fairness definitions can conflict when base rates differ. There is no single universal metric that settles every application. Document the intended use, affected people, acceptable trade-offs, and a review process.

PUT THE IDEA INTO CODE

A small experiment you can run.

Groups a and b are synthetic placeholders. Two positive examples per group are far too few for a reliable real-world conclusion.

fairness-and-error-slices.py
rows = [("a", 1, 1), ("a", 1, 0), ("a", 0, 0),
        ("b", 1, 1), ("b", 1, 1), ("b", 0, 1)]
for group in ["a", "b"]:
    positives = [(y, p) for g, y, p in rows if g == group and y == 1]
    missed = sum(p == 0 for _, p in positives)
    print(group, {"missed_positives": missed, "actual_positives": len(positives),
                  "false_negative_rate": missed/len(positives) if positives else None})
Copy code

Save the file, open your terminal in that folder, and run python fairness-and-error-slices.py. Use python3 or py if required by your installation. Setup guide

What to expect

The toy false-negative rates are 0.5 for a and 0 for b, based on two positives each.

YOUR TURN

Add denominators to a report.

  1. Calculate each group’s false-positive count and negative count.
  2. Compare with its false-negative result.
  3. Write one sentence explaining the sample-size limitation.
Compare with a suggested solution

Group a has no false positives among one negative; group b has one among one negative. These rates are extremely unstable. Report the counts and request more representative evidence before making a deployment claim.

CHECK YOUR UNDERSTANDING

One idea to take with you.

Why might equal accuracy be insufficient?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

NIST AI Risk Management Framework