Train a tiny next-word model
Learn what next-token prediction does using a transparent bigram model.
- Count token transitions
- Generate from a learned distribution
- Explain limitations of short context
Predict what follows
A bigram language model estimates the next token using only the current token. Count observed transitions in training text, normalize the counts, and sample a continuation. Start and end markers make sentence boundaries explicit.
This is a genuine statistical language model, but it has very limited context. It cannot keep track of long-range meaning, verify factual claims, or reason like a modern assistant. Its simplicity makes the learning and sampling mechanics visible.
Generation follows a distribution
Sampling can produce different outputs from the same initial token. A fixed seed makes a demonstration repeatable. Greedy decoding always chooses the most frequent continuation and can become repetitive.
An unseen context has no learned transition. Choose a documented fallback, such as stopping. Smoothing can assign some probability to unseen events, but it changes the distribution and can produce less plausible sequences in a tiny dataset.
A small experiment you can run.
The model learns transition counts from four original synthetic sentences. It calls no external AI service and downloads no weights.
import random
from collections import Counter, defaultdict
sentences = ["ai helps people learn", "ai helps people build", "people learn python", "people build tools"]
transitions = defaultdict(Counter)
for sentence in sentences:
words = ["<start>"] + sentence.split() + ["<end>"]
for left, right in zip(words, words[1:]):
transitions[left][right] += 1
rng, current, output = random.Random(17), "<start>", []
for _ in range(12):
choices = transitions.get(current)
if not choices:
break
current = rng.choices(list(choices), weights=list(choices.values()))[0]
if current == "<end>":
break
output.append(current)
print(" ".join(output))
Save the file, open your terminal in that folder, and run python a-tiny-language-model.py. Use python3 or py if required by your installation. Setup guide
The generated sentence uses learned transitions and stops at an end marker or the step limit.
Measure the effect of additional context.
- Find two sentences where the same word needs different continuations.
- Change the transition key to a pair of previous tokens.
- Explain why more context also requires more observations.
Compare with a suggested solution
A trigram model can distinguish two-word contexts, but each context appears less often. Sparse data creates more unseen combinations. Context length is a trade-off between useful conditioning and reliable estimation.
One idea to take with you.
Make it part of your progress.
Finish the practice and answer the knowledge check to mark this lesson complete.
Go deeper with primary documentation
Optional references for further study. This lesson and its examples were written for Artificials.
scikit-learn: text feature extractionAttention Is All You Need: original paper