Build an AI evaluation set you can trust
Collect representative cases, prevent leakage, and keep development examples separate from the evidence used to approve a release.

An evaluation set is a compact description of the work your AI system is expected to do. If the examples are too easy, too similar, or repeatedly used to tune the prompt, the resulting score can offer more reassurance than evidence. Good evaluation starts with the questions the data should help answer.
For a help-center assistant, those questions might be whether it finds the correct policy, recognizes missing evidence, and handles a request outside its supported scope. A single average should not obscure these different behaviors.
Collect ordinary cases and difficult cases deliberately
Start with representative requests obtained through an appropriate data-handling process. Remove unnecessary personal information and document where each example came from. Include ordinary traffic because it determines day-to-day usefulness, then add difficult cases to probe known failure modes.
Keep those groups identifiable. If you deliberately add many unusual adversarial requests, the combined pass rate no longer estimates the rate for ordinary production traffic. Both groups are valuable; they answer different questions. Report their results separately and explain how they were sampled.
Write the expected behavior before generating
Define the required facts, allowed alternatives, and unacceptable outcomes for each case. Some questions have several valid answers. Others should produce a clarification or an explicit statement that the evidence is missing. A reference answer can be a rubric rather than a single sentence to match.
For a support request about a feature that does not exist, the desired behavior might be to say it is unavailable and suggest a documented alternative. If your grader rewards mentioning the nonexistent feature, you have encoded the wrong product behavior into the test.
Keep development and holdout evidence separate
Use a development set to refine prompts and inspect failures. Reserve a holdout set for a later check. Repeatedly examining holdout mistakes and changing the system to fix them gradually turns that set into development data, even if no model weights were updated.
Scikit-learn's guidance on leakage explains why information from evaluation data must not influence training transformations. The wider lesson for an AI application is to record which examples influenced model, prompt, and retrieval decisions. A split in a spreadsheet is not enough if the same solutions enter the prompt elsewhere.
Split related material carefully
If several examples come from one support conversation or nearly identical document templates, a random row split can place close relatives on both sides. Group examples by the unit that needs to generalize, such as customer, document family, or time period.
For the help center, a time-based holdout can help examine behavior on recently added topics. It also changes the question being measured, so describe that choice. No split is universally correct; the split should reflect how the system will encounter unfamiliar work.
Version the evidence
Store a dataset version, rubric version, and a record of corrections. If a policy changes, update the expected answer deliberately and preserve the reason. Compare releases on a common version where possible so a changed score is not confused with a changed test.
Start small enough to inspect every case, then expand around observed gaps. Report counts as well as percentages, and avoid treating a tiny sample as a precise reliability estimate. The strongest evaluation set is one whose examples, judgments, and limits another reviewer can understand.