Learn AI/AI engineering
LESSON 27 / 36Intermediate 35 min with practice

Build an evaluation harness

Measure the whole workflow and turn failures into a development plan.

WHAT YOU WILL LEARN
  • Separate component and end-to-end metrics
  • Keep a fixed evaluation set
  • Report failures alongside averages

Evaluate the contract

An AI system can fail during input handling, retrieval, prediction, formatting, or tool execution. Measure the components separately, then check the final user outcome. A good retrieval score can coexist with poor answer quality.

Create cases with expected behavior, including ambiguous inputs, unsupported questions, and invalid requests. Keep development examples separate from a final evaluation set. Add newly discovered failures to a regression suite without pretending that repeatedly optimized cases are unseen.

Look behind the percentage

Report counts, not only percentages. Three correct results out of four provide much less evidence than 3,000 out of 4,000. Slice results by relevant conditions, such as language, input length, or document type, when enough examples exist.

A model judge can help review subjective outputs but is another fallible measurement tool. Specify the rubric, spot-check disagreements, and avoid giving an evaluator authority to invent ground truth. For deterministic outputs, direct checks are often better.

PUT THE IDEA INTO CODE

A small experiment you can run.

The harness reports the exact missing synonym instead of hiding it behind a 75% score. These four examples are a teaching fixture, not a deployment-quality evaluation.

evaluate-ai-systems.py
cases = [("reset password", "account"), ("invoice question", "billing"),
         ("unknown request", "review"), ("receipt missing", "billing")]
def route(text):
    return "account" if "password" in text else "billing" if "invoice" in text else "review"
failures = [(text, expected, route(text)) for text, expected in cases if route(text) != expected]
print({"correct": len(cases)-len(failures), "total": len(cases)})
for text, expected, actual in failures:
    print({"input": text, "expected": expected, "actual": actual})
Copy code

Save the file, open your terminal in that folder, and run python evaluate-ai-systems.py. Use python3 or py if required by your installation. Setup guide

What to expect

Three of four cases pass, and receipt missing is the failure.

YOUR TURN

Add meaningful evaluation slices.

  1. Tag cases as ordinary, synonym, or ambiguous.
  2. Calculate counts for each slice.
  3. Add a no-answer case and define the expected behavior before changing the rule.
Compare with a suggested solution

Keep a small table with input, expected output, and slice. Report correct/total for each slice and inspect failures. Small slices are diagnostic rather than strong statistical evidence.

CHECK YOUR UNDERSTANDING

One idea to take with you.

Why report the number of evaluated cases?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

Python functools module