Values, collections, and clear transformations
Represent examples using lists and dictionaries, then calculate a useful summary.
- Choose lists and dictionaries
- Transform values without hidden state
- Handle an empty collection
Represent the problem directly
A list preserves a sequence of examples. A dictionary names the attributes of an example. A set represents unique values, while a tuple is useful for a fixed grouping. Choosing a structure that mirrors the problem makes the code easier to reason about.
Keep types consistent. The text "12" is not the integer 12. Convert and validate values at the input boundary, rather than letting silent coercion spread through a training pipeline. Give intermediate results names that explain their meaning.
Make transformations inspectable
A list comprehension describes a transformation or filter in one place. Keep it short enough to read. When several decisions are involved, a regular loop with named steps is clearer.
Edge cases belong in the design: an empty list has no mean, a missing dictionary key is not automatically zero, and a category can be absent from a small batch. Represent an unavailable statistic explicitly instead of manufacturing a value.
A small experiment you can run.
The rows stay unchanged. Separate summaries make it easy to inspect the label set and text-length distribution before modeling.
examples = [{"text": "learn python", "label": "coding"},
{"text": "train a model", "label": "ml"},
{"text": "write a function", "label": "coding"}]
lengths = [len(row["text"].split()) for row in examples]
labels = sorted({row["label"] for row in examples})
mean_length = sum(lengths) / len(lengths) if lengths else None
print("Labels:", labels)
print("Mean words:", mean_length)
Save the file, open your terminal in that folder, and run python python-values-and-collections.py. Use python3 or py if required by your installation. Setup guide
The original labels are coding and ml, and the average length is 8 / 3 words.
Build a small dataset summary.
- Add a fourth example with a new label.
- Count examples per label using a dictionary.
- Repeat with an empty list and return an explicit unavailable mean.
Compare with a suggested solution
Initialize counts as an empty dictionary and increment counts[label] using counts.get(label, 0) + 1. Preserve None for the empty mean. Zero would misleadingly describe a measured average.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
Python language tutorialPython JSON module