Files, JSON, and dataset audits
Load structured data and catch invalid rows before training.
- Read JSON records
- Audit values and categories
- Keep raw data separate from cleaned data
Input data is a boundary
JSON represents objects, lists, strings, numbers, booleans, and null. A successful parse only proves the syntax is valid. You still need to check the shape, required keys, allowed labels, and types your application expects.
Keep the original input unchanged and write cleaned records separately. Store reasons for rejected rows so you can investigate systematic problems. Silently removing difficult examples may make evaluation look better while making the product worse.
Audit before optimizing
Count missing values, duplicates, category frequencies, and unusually long or short inputs. These simple checks often reveal mistakes more quickly than a model does. Deduplication rules should match the task: identical text with conflicting labels needs investigation.
Do not log private text simply because debugging is convenient. For a real dataset, choose the minimum necessary fields and access controls. The exercise uses invented records so it can be inspected safely without external downloads.
A small experiment you can run.
For a local file, replace json.loads(raw) with json.load(handle) inside with open("data.json", encoding="utf-8") as handle. The validation still applies.
import json
raw = '[{"text":"Reset password","label":"account"},{"text":" ","label":"billing"},{"text":"Receipt","label":"billing"}]'
rows = json.loads(raw)
valid, rejected = [], []
for index, row in enumerate(rows):
if (isinstance(row, dict) and isinstance(row.get("text"), str)
and row["text"].strip() and row.get("label") in ("account", "billing")):
valid.append({"text": row["text"].strip(), "label": row["label"]})
else:
rejected.append(index)
print(json.dumps({"accepted": len(valid), "rejected_indices": rejected}))
Save the file, open your terminal in that folder, and run python files-json-and-dataset-audits.py. Use python3 or py if required by your installation. Setup guide
The initial audit accepts two rows and rejects index 1.
Detect contradictory labels.
- Add two identical messages with different labels.
- Build a mapping from normalized text to its set of labels.
- Flag entries with more than one label for review.
Compare with a suggested solution
Use a separate normalized-text key, such as text.casefold().strip(), while preserving the original. A set of two labels for one key is a conflict, not a reason to choose the first label automatically.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
Python language tutorialPython JSON module