Project: a support-message classifier
Train a tiny text classifier, evaluate held-out messages, and inspect its errors.
- Train multinomial Naive Bayes
- Keep vocabulary inside training
- Evaluate a complete prediction path
Build a complete baseline
This project uses multinomial Naive Bayes. Each class has a prior based on its training frequency and a distribution over vocabulary words. Add-one smoothing keeps observed-vocabulary words from getting zero probability in a class.
For a query, sum the log prior and word log probabilities, then choose the highest score. Summing logs avoids underflow from multiplying many small numbers. The conditional-independence assumption is a simplification; correlated words are treated as separate evidence.
Evaluate honestly
The dataset is deliberately small and invented. It demonstrates a full training-to-inference path, not a production support classifier. Unknown words are ignored, so a query with no known terms falls back to review.
The selected class score is not a calibrated probability. Inspect missed synonyms, mixed-topic messages, negation, and unknown language. Expand the labeling guide and evaluation set before adding complexity. Keep any final test examples out of vocabulary construction and tuning.
A small experiment you can run.
The script learns word counts only from the four training messages. The refund example intentionally reveals a missing concept in the vocabulary.
import math, re
from collections import Counter, defaultdict
train = [("reset password login", "account"), ("account sign in password", "account"),
("invoice payment billing", "billing"), ("receipt charge payment", "billing")]
def tokens(text):
return re.findall(r"[a-z]+", text.lower())
counts, classes = defaultdict(Counter), Counter()
for text, label in train:
counts[label].update(tokens(text))
classes[label] += 1
vocab = set().union(*(set(c) for c in counts.values()))
def predict(text):
words = [w for w in tokens(text) if w in vocab]
if not words:
return "review"
scores = {}
for label in classes:
total = sum(counts[label].values()) + len(vocab)
scores[label] = math.log(classes[label]/len(train)) + sum(
math.log((counts[label][word]+1)/total) for word in words)
return max(sorted(scores), key=scores.get)
test = [("forgot login password", "account"), ("payment invoice question", "billing"),
("galactic weather", "review"), ("refund request", "billing")]
correct = 0
for text, expected in test:
actual = predict(text)
correct += actual == expected
print({"text": text, "expected": expected, "prediction": actual})
print("Correct:", correct, "of", len(test))
Save the file, open your terminal in that folder, and run python project-support-message-classifier.py. Use python3 or py if required by your installation. Setup guide
The initial script gets 3 of 4 held-out fixture cases correct; refund request is sent to review.
Extend the classifier without contaminating its evaluation.
- Create a separate development set with mixed topics and unfamiliar wording.
- Add original training messages for missing concepts, then freeze the vocabulary.
- Evaluate once on a new held-out set and write a short model card with counts and failures.
Compare with a suggested solution
Use development errors to improve training coverage. Do not report the repeatedly inspected four-case fixture as an unbiased final score. Keep no-known-word queries in review and consider an ambiguity rule validated on separate examples.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
Python SQLite modulePython HTTP server: limitations