Create a repeatable train/test split
Keep evaluation data separate and make the boundary easy to inspect.
- Shuffle without changing the original
- Separate development and final evaluation
- Check that partitions do not overlap
Separate the jobs of your data
Training data fits parameters. Validation data helps choose settings. Test data estimates performance after those choices are complete. If you repeatedly inspect the test score and adapt the model, that set has effectively become validation data.
A random split can work when examples are independent and identically distributed. Time series and grouped entities need different splits. Small or rare classes may require stratification so that each partition contains enough examples to evaluate.
Record the boundary
Keep row identifiers as well as values. Saving split IDs makes it possible to compare two models on exactly the same test cases. A fixed seed is helpful, but a later change in the input ordering can still change the split.
Fit preprocessing only on the training partition. A vocabulary, mean, standard deviation, or feature selection rule can leak information even when the prediction model itself never sees the test labels.
A small experiment you can run.
This example partitions identifiers, which can be saved alongside a dataset version. It does not assume that a random split is appropriate for every problem.
import random
ids = list(range(30))
shuffled = ids.copy()
random.Random(17).shuffle(shuffled)
train, validation, test = shuffled[:18], shuffled[18:24], shuffled[24:]
assert len(set(train + validation + test)) == len(ids)
assert set(train).isdisjoint(test)
print("Train:", len(train), "Validation:", len(validation), "Test:", len(test))
print("Test IDs:", sorted(test))
Save the file, open your terminal in that folder, and run python reproducible-data-splits.py. Use python3 or py if required by your installation. Setup guide
The random example creates 18 training, 6 validation, and 6 test IDs.
Compare random and time-based splitting.
- Create 30 observations with increasing timestamps.
- Reserve the latest six as a temporal test set.
- Explain which split matches next-month forecasting.
Compare with a suggested solution
Use chronological order for the temporal split. Random shuffling mixes future and past patterns, which can hide a change in conditions. For forecasting, keep the latest period unseen and avoid features computed from later observations.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
Python language tutorialPython JSON module