Learn AI/Responsible AI
LESSON 32 / 36Intermediate 30 min with practice

Monitor changes and write a model card

Make a system’s intended use and limits inspectable after release.

WHAT YOU WILL LEARN
  • Document the evaluation setting
  • Distinguish input drift from performance loss
  • Plan rollback and review

Document what was actually built

A model card describes purpose, supported inputs, training and evaluation data, metrics, limitations, and inappropriate uses. Include the exact model and data versions, evaluation date, and known failure examples.

Do not substitute broad claims such as “unbiased” or “safe” for evidence. State what was measured, how many examples were used, and what remains uncertain. A small educational model should be labeled as such.

Watch the system over time

Input drift means the distribution of inputs changes. It may signal a problem but does not prove the error rate increased. When labels arrive, measure actual outcomes as well as proxy indicators such as input length, unknown vocabulary, latency, and abstention rate.

Choose alert thresholds based on historical variation and operational needs. Define who investigates an alert and how to roll back a model or disable an action. Keep a previous working version and the configuration that produced it.

PUT THE IDEA INTO CODE

A small experiment you can run.

This card reports exact fixture counts and limitations. It is not a certification or evidence that the model is ready for public decision-making.

monitoring-and-model-cards.py
import json
card = {"name": "Study-topic demo", "version": "1.0", "purpose": "Educational routing experiment",
        "data": "Original synthetic examples", "evaluation": {"correct": 3, "total": 4},
        "limitations": ["Tiny English-only fixture", "No real-user validation"],
        "not_for": ["Consequential automated decisions"], "fallback": "Human review"}
print(json.dumps(card, indent=2))
Copy code

Save the file, open your terminal in that folder, and run python monitoring-and-model-cards.py. Use python3 or py if required by your installation. Setup guide

What to expect

The output is a structured card containing intended use, evidence, and limitations.

YOUR TURN

Write a model card for your first project.

  1. Record the dataset and split IDs.
  2. Include two concrete failure examples.
  3. Describe one rollback trigger and one condition that requires more evaluation.
Compare with a suggested solution

A useful trigger might be a sustained increase in invalid outputs, while a new supported language requires new evaluation. Distinguish these operational policies from claims about model quality.

CHECK YOUR UNDERSTANDING

One idea to take with you.

Does input drift alone prove model accuracy decreased?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

NIST AI Risk Management Framework