Learn AI/Python
LESSON 07 / 36Beginner 30 min with practice

Files, JSON, and dataset audits

Load structured data and catch invalid rows before training.

WHAT YOU WILL LEARN
  • Read JSON records
  • Audit values and categories
  • Keep raw data separate from cleaned data

Input data is a boundary

JSON represents objects, lists, strings, numbers, booleans, and null. A successful parse only proves the syntax is valid. You still need to check the shape, required keys, allowed labels, and types your application expects.

Keep the original input unchanged and write cleaned records separately. Store reasons for rejected rows so you can investigate systematic problems. Silently removing difficult examples may make evaluation look better while making the product worse.

Audit before optimizing

Count missing values, duplicates, category frequencies, and unusually long or short inputs. These simple checks often reveal mistakes more quickly than a model does. Deduplication rules should match the task: identical text with conflicting labels needs investigation.

Do not log private text simply because debugging is convenient. For a real dataset, choose the minimum necessary fields and access controls. The exercise uses invented records so it can be inspected safely without external downloads.

PUT THE IDEA INTO CODE

A small experiment you can run.

For a local file, replace json.loads(raw) with json.load(handle) inside with open("data.json", encoding="utf-8") as handle. The validation still applies.

files-json-and-dataset-audits.py
import json
raw = '[{"text":"Reset password","label":"account"},{"text":" ","label":"billing"},{"text":"Receipt","label":"billing"}]'
rows = json.loads(raw)
valid, rejected = [], []
for index, row in enumerate(rows):
    if (isinstance(row, dict) and isinstance(row.get("text"), str)
            and row["text"].strip() and row.get("label") in ("account", "billing")):
        valid.append({"text": row["text"].strip(), "label": row["label"]})
    else:
        rejected.append(index)
print(json.dumps({"accepted": len(valid), "rejected_indices": rejected}))
Copy code

Save the file, open your terminal in that folder, and run python files-json-and-dataset-audits.py. Use python3 or py if required by your installation. Setup guide

What to expect

The initial audit accepts two rows and rejects index 1.

YOUR TURN

Detect contradictory labels.

  1. Add two identical messages with different labels.
  2. Build a mapping from normalized text to its set of labels.
  3. Flag entries with more than one label for review.
Compare with a suggested solution

Use a separate normalized-text key, such as text.casefold().strip(), while preserving the original. A set of two labels for one key is a conflict, not a reason to choose the first label automatically.

CHECK YOUR UNDERSTANDING

One idea to take with you.

What does valid JSON guarantee?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

Python language tutorialPython JSON module