Tokenization and vocabularies
Turn text into model inputs while keeping the representation’s limits visible.
- Build a training vocabulary
- Handle unknown words
- Distinguish toy word splitting from production tokenizers
Tokens are representation units
A tokenizer maps text into units a model can process. Those units may be words, subwords, characters, or bytes. Token counts therefore depend on the tokenizer, language, and input; character count is not a reliable universal conversion.
Our learning tokenizer lowercases and extracts basic Latin words with a regular expression. It deliberately simplifies punctuation and multilingual text. Real applications need a tokenizer that matches the exact model and handles the intended languages.
Fit the vocabulary at the right time
Build a learned vocabulary from training text only. Reserve an unknown token for inputs absent from the vocabulary. If you rebuild the mapping on each request, the same ID can change meaning.
Save the vocabulary with the model. Decide how case, punctuation, whitespace, and empty text are handled. Some normalization can erase information the task needs, such as case in code or punctuation in sentiment.
A small experiment you can run.
Zero is reserved for unknown words. The word new is absent from training and therefore maps to zero.
import re
training = ["Learn AI with Python", "Build useful AI tools"]
def tokenize(text):
return re.findall(r"[a-z]+", text.casefold())
words = sorted({word for text in training for word in tokenize(text)})
vocabulary = {word: i+1 for i, word in enumerate(words)}
encoded = [vocabulary.get(word, 0) for word in tokenize("Learn new tools")]
print("Vocabulary:", vocabulary)
print("Token IDs:", encoded)
Save the file, open your terminal in that folder, and run python tokenization-and-vocabularies.py. Use python3 or py if required by your installation. Setup guide
The encoded sequence contains zero for new.
Expose the tokenizer’s limitations.
- Try a phrase containing accented characters.
- Try source code with punctuation.
- Document which distinctions the regular expression discards.
Compare with a suggested solution
The pattern only includes a-z, so it can split or discard other scripts and accented letters. Code punctuation also disappears. Keep this tokenizer labeled as a teaching simplification rather than using it uncritically in a multilingual product.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
scikit-learn: text feature extractionAttention Is All You Need: original paper