An animated journey, from your keyboard to the silicon and back
How AI works
Follow one question down through every layer of a modern AI system: the words, the mathematics, the learning, the chips, the data centres. Then back up to the answer.
The question we follow: Why is the sky blue?
Scroll to go inChapter 1 of 22
Words become pieces
A model cannot read letters. First, your question is cut into tokens.
You type a question: “Why is the sky blue?” To a computer it is a string of characters. A language model works with a fixed set of pieces instead, so the first job is to find them.
Cut into tokens
A tokenizer splits the text into tokens: common words whole, rarer words in parts, punctuation on its own. The space before a word usually travels with it, which is why “ sky” and “sky” can be different tokens.
In English, one token averages about four characters. OpenAI
Look each one up
Every token has a number: its place in the model’s vocabulary, a list fixed before training begins. Vocabularies hold from tens of thousands to a few hundred thousand tokens. The numbers here are illustrative; each tokenizer has its own.
A row of numbers
Your question is now a row of whole numbers. From here on, everything the model does to it is arithmetic.
Chapter 2 of 22
Numbers with meaning
Each token becomes a long list of numbers: a point in a space where related words sit close together.
A token’s number is only a label. The model swaps each one for a vector: a list of numbers it learned during training. Our six tokens become six columns of values.
The largest Llama 3 model gives each token 16,384 numbers. Meta, 2024
A space of meaning
Read as coordinates, each vector is a point. Training arranges the points so that words used in similar ways end up near each other. The picture squeezes thousands of dimensions into three.
Neighbours
“sky” lands among cloud, air and rain; “blue” among the other colours. Nobody wrote those rules down: they come from how words are used across the training text.
Directions mean something too
In a well-known early result, the step from “man” to “woman” matched the step from “king” to “queen”. Directions in the space can carry ideas like these. The layout here is illustrative.
Word vectors: Mikolov et al., 2013
Chapter 3 of 22
One artificial neuron
The basic part is simple: multiply, add, bend.
A neuron takes a few numbers in and sends one number out. This one has four inputs.
Weigh each input
Each input is multiplied by a weight. A large weight makes its input count for a lot; a negative weight pushes the other way. The weights are what training adjusts.
Add, then bend
The neuron adds the products and a bias, then passes the sum through an activation function: a bend that lets a network describe more than straight lines. GELU, drawn here, is common in language models.
1.32 − 0.20 = 1.12, and GELU(1.12) ≈ 0.97. Hendrycks and Gimpel, 2016
One of billions
A model is built from enormous numbers of these simple parts. Its parameters, the weights and biases, are counted in billions.
Llama 3.1’s largest model has 405 billion parameters. Meta
Chapter 4 of 22
Layer upon layer
Neurons are arranged in layers, and each layer feeds the next.
The token vectors enter the first layer.
A wave through the network
Every neuron in a layer reads the outputs of the layer before, weighs them, adds them and bends the sum. The signal sweeps forward, one layer at a time.
Scores come out
At the far end, the last layer produces the numbers the model will choose from.
Deep learning
Stacking many layers is what makes a network deep. Language models are much deeper than this sketch, and their layers are of a particular kind: transformer blocks.
Chapter 5 of 22
Attention
Each word looks back at the words before it to work out what it means here.
On its own, “blue” could be a colour or a mood. To know which, the model lets each position gather information from the positions before it.
Weighing the context
“blue” scores every earlier word by how relevant it is, then takes a weighted mix of what those words carry. Here it leans most on “sky”.
Many heads
An attention layer runs several of these side by side, each with its own weights. One head may follow what is being described, another the word just before, another the question being asked.
The pattern
As a grid, each row is one position’s attention, adding up to one. The upper half is empty: in a language model a word can look back, never ahead.
Attention is the heart of the transformer. Vaswani et al., 2017
Chapter 6 of 22
The transformer
The same block, repeated: attention, then a small network, over and over.
A transformer block has two parts. Attention lets positions exchange information; a feed-forward network then works on each position by itself.
Stacked
Blocks are stacked, each refining what the block below produced.
The residual stream
Through the middle runs a stream of numbers for every position. Each block reads from it and adds its own contribution back, so information climbs without being lost.
Out of the top
After the last block, the stream at the final position becomes a score for every token in the vocabulary.
Llama 3.1 405B stacks 126 of these blocks. Meta, 2024
Chapter 7 of 22
One token at a time
A model does not write an answer. It picks one token, then the next.
The tower gives a score to every token in the vocabulary: tens of thousands of numbers.
Scores into chances
Softmax turns the scores into probabilities that add up to one. Temperature sets how bold the choice is: low temperature sharpens the top candidates, high temperature evens them out. Try the slider.
Draw one
A token is drawn at random, weighted by those probabilities. Usually a likely token wins; sometimes a less likely one does, which is why the same question can get different answers.
And again
The chosen token joins the text, and the whole model runs again to choose the next. The answer grows a token at a time until the model chooses to stop.
Part two: learning
But where do all those weights come from?
Chapter 8 of 22
Learning by going downhill
Training adjusts the weights, a little at a time, to make fewer mistakes.
Training begins with random weights, so the model’s predictions are poor. Loss measures how poor; here, it is the height of the landscape.
Which way is down?
For every weight, calculus gives the gradient: the direction in which the loss rises fastest. Training takes a small step the opposite way.
Step after step
Repeating the step walks the weights downhill. Not every walk ends in the deepest valley: this one does, while a second walker settles in a shallow dip.
At scale
A real model walks through billions of dimensions at once, not two, one batch of training text after another.
Llama 3 was pretrained on more than 15 trillion tokens. Meta
Chapter 9 of 22
Blame, passed backwards
Backpropagation works out how much each weight contributed to the error.
The model’s prediction is compared with the right answer: the token that really came next in the training text.
The error flows back
Starting at the output, the error is passed backwards through the network, layer by layer, with the chain rule of calculus.
Every weight moves
Each weight receives its share of the blame and moves a little in the direction that would have made the error smaller.
Forward, backward, repeat
Training alternates the two passes, again and again, across all its text.
Backpropagation: Rumelhart, Hinton and Williams, 1986
Chapter 10 of 22
From predictor to assistant
Training on web text makes a model that continues text. A second stage teaches it to answer.
A model trained only to predict the next token of web text is a base model. Asked a question, it may simply carry on the page: more questions, a list, a heading.
Shown good answers
People write example answers to many prompts, and the model is trained on them the same way as before, one token at a time. This is supervised fine-tuning.
Asked which is better
The model writes several answers and people rank them. A second model, the reward model, learns to predict which answers people prefer.
Nudged towards preferred answers
Reinforcement learning then adjusts the model to earn higher scores from the reward model. Newer methods, such as direct preference optimisation, learn from the rankings without a separate reward model.
People preferred answers from a 1.3-billion-parameter tuned model to those of the 175-billion-parameter GPT-3: Ouyang et al., 2022. Direct preference optimisation: Rafailov et al., 2023
Part three: hardware
All of it is arithmetic. An astonishing amount of it.
Chapter 11 of 22
It is all multiplication
Under every layer is one operation, repeated enormously: matrix multiplication.
Most of a layer’s work is one calculation: a matrix of inputs multiplied by a matrix of weights.
One cell at a time
Each output cell pairs a row of inputs with a column of weights, multiplies them pair by pair and adds the results. The first cell here takes six multiplications.
Every cell at once
No cell depends on another, so they can all be worked out at the same moment. Hardware for AI is built around that independence.
At real size
A real layer multiplies matrices with thousands of rows and columns. Writing one token takes roughly two operations for every parameter in the model.
For a 70-billion-parameter model, that is about 140 billion operations per token. Kaplan et al., 2020
Chapter 12 of 22
Why graphics chips
A job made of billions of identical sums needs many hands, not a few clever ones.
A CPU has a handful of powerful cores, built to race through complicated tasks one after another.
A few at a time
Given thousands of independent sums, a CPU works through them a few at a time.
Thousands at once
A GPU has thousands of simpler cores that run the same instruction on different numbers at once. Designed to colour the pixels of games, the idea suits neural networks just as well.
Tensor cores
Modern GPUs add tensor cores: units that multiply small matrices in a single step, usually at reduced precision.
An NVIDIA H100 (SXM) has 16,896 CUDA cores and 528 Tensor Cores. NVIDIA
Chapter 13 of 22
The real bottleneck
Arithmetic is fast. Fetching the numbers to do it on is slow.
Around the compute die sit stacks of high-bandwidth memory, HBM, which hold the model’s weights.
Every weight, every token
To write one token, the chip reads every weight from memory. While the data streams in, the arithmetic units mostly wait.
The arithmetic of waiting
A 70-billion-parameter model at 16 bits is 140 GB. At 3.35 TB/s, one full read takes about 42 thousandths of a second, which caps a single conversation near 24 tokens a second, even if it fitted on one chip.
H100 (SXM): 80 GB of HBM3 at 3.35 TB/s. NVIDIA
Share each read
Serving many requests together lets each read of the weights do many people’s work. That is why AI services batch requests.
Part four: scale
One GPU is not enough. Not even close.
Chapter 14 of 22
Why so many GPUs
Large models outgrow one GPU’s memory, and training needs far more still.
A 70-billion-parameter model at 16 bits needs 140 GB for its weights alone. One GPU holds 80 GB. Try the other sizes below the picture.
Split it
So the model is split across GPUs. Each holds a slice of every layer, and they trade results over fast links after each one.
Training needs more
Training also keeps gradients and optimizer state: about 16 bytes per parameter with the common Adam optimizer, before counting the activations.
16 bytes per parameter for mixed-precision Adam. Rajbhandari et al., 2019
Keep in step
To learn from more text at once, copies of the model train on different batches, then combine their gradients so every copy takes the same step.
Chapter 15 of 22
The training bill
Large models take millions of GPU hours. The only way through is many GPUs at once.
Meta reports that training its largest Llama 3.1 model took 30.84 million GPU hours.
Llama 3.1 405B: 30.84M H100 GPU hours. Meta
On one GPU
Done on a single GPU, that would take about 3,500 years.
On sixteen thousand
Meta trained it on up to 16,384 GPUs at once. Divided perfectly, that is about eleven weeks; real training also loses time to failures and waiting. Move the slider to see other counts.
Up to 16,384 H100 GPUs. Meta, 2024
And the power
Each of those GPUs can draw up to 700 watts: together about 11.5 megawatts, before cooling, networking and the rest of the building.
H100 (SXM): up to 700 W. NVIDIA
Chapter 16 of 22
A team of experts
Some of the largest models use only a small part of themselves for each token.
In the transformer, every token passes through every weight of each feed-forward block. A bigger block means more arithmetic for every token.
A router picks
A mixture-of-experts layer splits that block into several smaller ones, the experts, and adds a small router. For each token the router scores the experts and sends the token to the best few, blending their answers. Mixtral uses eight experts in each layer and picks two.
Mixtral: Jiang et al., 2024
Each token, its own pair
Different tokens take different experts; the router learns its choices during training. They are not tidy subjects: in Mixtral, experts did not split by topic, and the patterns found were closer to grammar.
All in memory, a few at work
Every expert must still sit in GPU memory, but each token only does the arithmetic of the ones it visits. A model can hold far more than it uses on any one token.
Mixtral 8x7B: 47B parameters, 13B used for each token. DeepSeek-V3: 671B parameters, 37B used for each token. Mixtral, DeepSeek-V3
Chapter 17 of 22
Inside the data centre
Zoom out, from one chip to the building it lives in.
One GPU: a compute die, with memory stacked beside it.
A server
Eight GPUs share a server, joined by fast links, such as NVIDIA’s NVLink, so that they can work almost as one.
Racks and networks
Servers fill racks, and racks are wired together by a network that carries results and gradients between thousands of GPUs.
Power in, heat out
Almost all the electricity a data centre uses ends up as heat. Cold air or liquid flows in, hot air or liquid flows out, around the clock.
Part five: back to you
Now, back up to your question.
Chapter 18 of 22
Serving your question
When you press Enter, your question joins many others on a GPU somewhere.
Your question arrives as tokens.
Prefill
The model reads the whole prompt in one pass, every position in parallel. This is the pause before the first word appears.
Decode
Then it writes one token per step. Each step reuses what earlier tokens left in the KV cache, so nothing is computed twice, and the cache grows with the conversation.
Many at once
A server interleaves many conversations and adds new ones as others finish, so every read of the weights serves as many people as it can.
Chapter 19 of 22
Smaller numbers
Store each weight in fewer bits, and the model shrinks, at a small cost in accuracy.
Training usually keeps a 32-bit copy of every weight: precise, and large. For a 70-billion-parameter model, 280 GB.
16 bits
Most models run in 16 bits. Here the weight moves by less than a ten-thousandth.
8 bits
At 8 bits, the values a weight can take thin out. This one lands about 0.02 away, and the model halves again.
4 bits
At 4 bits, sixteen levels times a shared scale stand in for every weight. The model fits in 35 GB, on one GPU, at some cost in quality.
Chapter 20 of 22
Beyond words
The same ideas read pictures, and draw them.
Vision models cut an image into patches: small squares of pixels.
Patches as tokens
Each patch becomes a vector, like a word, and a transformer reads them together.
“An image is worth 16×16 words.” Dosovitskiy et al., 2020
Drawing from noise
Image generators work the other way round. Training teaches a model to remove a little noise from pictures, one step at a time.
Step by step
To draw, it starts from pure noise and removes it step after step, guided by your words, until a picture is left. The steps shown are one picture’s noised versions, played backwards.
Diffusion models: Ho, Jain and Abbeel, 2020
Chapter 21 of 22
Models that use tools
Given tools, a model can look things up and act, not only write.
Your question reaches the model.
Calling a tool
When it lacks something, the model can write a request for a tool, such as a search, instead of an answer. The software around it decides whether that call is allowed.
Results join the context
The tool’s result is added to the model’s context, the same way your question was.
An answer, with its source
With the result in view, the model writes its answer, and can point to where it came from.
Chapter 22 of 22
What it is not
A model predicts likely words. Likely is not the same as true.
An answer can read fluently and confidently.
Likely, not true
Each token was chosen because it was likely in context. Nothing in that process checks the claim against the world.
Check a source
A reliable source settles it: sunlight scatters off the gases in the air, and blue light scatters more than red.
Why is the sky blue? NASA Space Place
The last check is yours
Tools and sources help, but judgment stays with the reader. Ask where a claim comes from before you rely on it.
The way back up
That is how AI works.
You asked: Why is the sky blue?
Air molecules scatter blue light more than red.
Tokens, vectors and attention; a walk downhill through billions of weights; arithmetic on thousands of cores, in buildings full of them; and back to one sentence on your screen.
Sources
The pictures are simplified teaching models drawn in your browser, not a live view of any product. Figures and results come from these sources; illustrative values are marked where they appear.
- OpenAI: What are tokens and how to count them?
- Mikolov et al., 2013: Efficient Estimation of Word Representations in Vector Space
- Hendrycks and Gimpel, 2016: Gaussian Error Linear Units (GELUs)
- Vaswani et al., 2017: Attention Is All You Need
- Rumelhart, Hinton and Williams, 1986: Learning representations by back-propagating errors
- Ouyang et al., 2022: Training language models to follow instructions with human feedback
- Rafailov et al., 2023: Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Kaplan et al., 2020: Scaling Laws for Neural Language Models
- Rajbhandari et al., 2019: ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
- Meta, 2024: The Llama 3 Herd of Models
- Meta: Llama 3.1 model card
- Meta, 2024: Introducing Meta Llama 3
- Jiang et al., 2024: Mixtral of Experts
- DeepSeek-AI, 2024: DeepSeek-V3 Technical Report
- NVIDIA: H100 Tensor Core GPU specifications
- NVIDIA: Hopper architecture whitepaper
- Dosovitskiy et al., 2020: An Image is Worth 16x16 Words
- Ho, Jain and Abbeel, 2020: Denoising Diffusion Probabilistic Models
- NASA Space Place: Why Is the Sky Blue?