Data, labels, and the unit of prediction
Understand what one example represents and why repeated examples distort evaluation.
- Identify features and targets
- Recognize label ambiguity
- Separate entities across splits
One row needs a meaning
A dataset is not just a table: each row represents a specific observation. If a row is a support message, the target might be the team that should answer it. If the target instead records the team that happened to answer, the labels may capture staffing decisions rather than the desired behavior.
Write a labeling guide with positive examples, borderline cases, and an unknown option. Check disagreement between annotators. More rows with inconsistent targets do not necessarily improve a model. Keep a record of collection permissions and the population represented.
Independence matters
Two messages from the same conversation are related. Putting one in training and the other in testing can make the task artificially easy. Similar problems arise with duplicate documents, repeated customers, and multiple images of the same object.
Split by the entity or time boundary that matches deployment. A future-facing forecast needs later observations in its test set. A model intended for new customers should be evaluated on customers absent from training.
A small experiment you can run.
The split keeps an entire conversation together. It demonstrates the boundary, not a statistically meaningful evaluation on four examples.
rows = [
{"conversation": "a", "text": "reset password", "label": "account"},
{"conversation": "a", "text": "still cannot sign in", "label": "account"},
{"conversation": "b", "text": "invoice missing", "label": "billing"},
{"conversation": "c", "text": "download receipt", "label": "billing"},
]
train = [r for r in rows if r["conversation"] != "c"]
test = [r for r in rows if r["conversation"] == "c"]
assert {r["conversation"] for r in train}.isdisjoint(r["conversation"] for r in test)
print(len(train), len(test))
Save the file, open your terminal in that folder, and run python data-and-labels.py. Use python3 or py if required by your installation. Setup guide
The example prints 3 1 and passes the independence assertion.
Design a split for a document collection.
- Identify the entity that could create related rows.
- List a grouping key for duplicates or document families.
- Decide whether deployment is about future documents or unseen document families.
Compare with a suggested solution
Group versions of the same source together. If future performance is the goal, reserve a later collection period and ensure copied passages do not cross the boundary. Preserve the final test set until model choices are complete.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
Python: virtual environments