An animated journey, from your keyboard to the silicon and back

How AI works

Follow one question down through every layer of a modern AI system: the words, the mathematics, the learning, the chips, the data centres. Then back up to the answer.

The question we follow: Why is the sky blue?

Scroll to go in

Chapter 1 of 22

Words become pieces

A model cannot read letters. First, your question is cut into tokens.

You type a question: “Why is the sky blue?” To a computer it is a string of characters. A language model works with a fixed set of pieces instead, so the first job is to find them.

Cut into tokens

A tokenizer splits the text into tokens: common words whole, rarer words in parts, punctuation on its own. The space before a word usually travels with it, which is why “ sky” and “sky” can be different tokens.

In English, one token averages about four characters. OpenAI

Look each one up

Every token has a number: its place in the model’s vocabulary, a list fixed before training begins. Vocabularies hold from tens of thousands to a few hundred thousand tokens. The numbers here are illustrative; each tokenizer has its own.

A row of numbers

Your question is now a row of whole numbers. From here on, everything the model does to it is arithmetic.

Chapter 2 of 22

Numbers with meaning

Each token becomes a long list of numbers: a point in a space where related words sit close together.

A token’s number is only a label. The model swaps each one for a vector: a list of numbers it learned during training. Our six tokens become six columns of values.

The largest Llama 3 model gives each token 16,384 numbers. Meta, 2024

A space of meaning

Read as coordinates, each vector is a point. Training arranges the points so that words used in similar ways end up near each other. The picture squeezes thousands of dimensions into three.

Neighbours

“sky” lands among cloud, air and rain; “blue” among the other colours. Nobody wrote those rules down: they come from how words are used across the training text.

Directions mean something too

In a well-known early result, the step from “man” to “woman” matched the step from “king” to “queen”. Directions in the space can carry ideas like these. The layout here is illustrative.

Word vectors: Mikolov et al., 2013

Chapter 3 of 22

One artificial neuron

The basic part is simple: multiply, add, bend.

A neuron takes a few numbers in and sends one number out. This one has four inputs.

Weigh each input

Each input is multiplied by a weight. A large weight makes its input count for a lot; a negative weight pushes the other way. The weights are what training adjusts.

Add, then bend

The neuron adds the products and a bias, then passes the sum through an activation function: a bend that lets a network describe more than straight lines. GELU, drawn here, is common in language models.

1.32 − 0.20 = 1.12, and GELU(1.12) ≈ 0.97. Hendrycks and Gimpel, 2016

One of billions

A model is built from enormous numbers of these simple parts. Its parameters, the weights and biases, are counted in billions.

Llama 3.1’s largest model has 405 billion parameters. Meta

Chapter 4 of 22

Layer upon layer

Neurons are arranged in layers, and each layer feeds the next.

The token vectors enter the first layer.

A wave through the network

Every neuron in a layer reads the outputs of the layer before, weighs them, adds them and bends the sum. The signal sweeps forward, one layer at a time.

Scores come out

At the far end, the last layer produces the numbers the model will choose from.

Deep learning

Stacking many layers is what makes a network deep. Language models are much deeper than this sketch, and their layers are of a particular kind: transformer blocks.

Chapter 5 of 22

Attention

Each word looks back at the words before it to work out what it means here.

On its own, “blue” could be a colour or a mood. To know which, the model lets each position gather information from the positions before it.

Weighing the context

“blue” scores every earlier word by how relevant it is, then takes a weighted mix of what those words carry. Here it leans most on “sky”.

Many heads

An attention layer runs several of these side by side, each with its own weights. One head may follow what is being described, another the word just before, another the question being asked.

The pattern

As a grid, each row is one position’s attention, adding up to one. The upper half is empty: in a language model a word can look back, never ahead.

Attention is the heart of the transformer. Vaswani et al., 2017

Chapter 6 of 22

The transformer

The same block, repeated: attention, then a small network, over and over.

A transformer block has two parts. Attention lets positions exchange information; a feed-forward network then works on each position by itself.

Stacked

Blocks are stacked, each refining what the block below produced.

The residual stream

Through the middle runs a stream of numbers for every position. Each block reads from it and adds its own contribution back, so information climbs without being lost.

Out of the top

After the last block, the stream at the final position becomes a score for every token in the vocabulary.

Llama 3.1 405B stacks 126 of these blocks. Meta, 2024

Chapter 7 of 22

One token at a time

A model does not write an answer. It picks one token, then the next.

The tower gives a score to every token in the vocabulary: tens of thousands of numbers.

Scores into chances

Softmax turns the scores into probabilities that add up to one. Temperature sets how bold the choice is: low temperature sharpens the top candidates, high temperature evens them out. Try the slider.

Draw one

A token is drawn at random, weighted by those probabilities. Usually a likely token wins; sometimes a less likely one does, which is why the same question can get different answers.

And again

The chosen token joins the text, and the whole model runs again to choose the next. The answer grows a token at a time until the model chooses to stop.

Part two: learning

But where do all those weights come from?

Chapter 8 of 22

Learning by going downhill

Training adjusts the weights, a little at a time, to make fewer mistakes.

Training begins with random weights, so the model’s predictions are poor. Loss measures how poor; here, it is the height of the landscape.

Which way is down?

For every weight, calculus gives the gradient: the direction in which the loss rises fastest. Training takes a small step the opposite way.

Step after step

Repeating the step walks the weights downhill. Not every walk ends in the deepest valley: this one does, while a second walker settles in a shallow dip.

At scale

A real model walks through billions of dimensions at once, not two, one batch of training text after another.

Llama 3 was pretrained on more than 15 trillion tokens. Meta

Chapter 9 of 22

Blame, passed backwards

Backpropagation works out how much each weight contributed to the error.

The model’s prediction is compared with the right answer: the token that really came next in the training text.

The error flows back

Starting at the output, the error is passed backwards through the network, layer by layer, with the chain rule of calculus.

Every weight moves

Each weight receives its share of the blame and moves a little in the direction that would have made the error smaller.

Forward, backward, repeat

Training alternates the two passes, again and again, across all its text.

Backpropagation: Rumelhart, Hinton and Williams, 1986

Chapter 10 of 22

From predictor to assistant

Training on web text makes a model that continues text. A second stage teaches it to answer.

A model trained only to predict the next token of web text is a base model. Asked a question, it may simply carry on the page: more questions, a list, a heading.

Shown good answers

People write example answers to many prompts, and the model is trained on them the same way as before, one token at a time. This is supervised fine-tuning.

Asked which is better

The model writes several answers and people rank them. A second model, the reward model, learns to predict which answers people prefer.

Nudged towards preferred answers

Reinforcement learning then adjusts the model to earn higher scores from the reward model. Newer methods, such as direct preference optimisation, learn from the rankings without a separate reward model.

People preferred answers from a 1.3-billion-parameter tuned model to those of the 175-billion-parameter GPT-3: Ouyang et al., 2022. Direct preference optimisation: Rafailov et al., 2023

Part three: hardware

All of it is arithmetic. An astonishing amount of it.

Chapter 11 of 22

It is all multiplication

Under every layer is one operation, repeated enormously: matrix multiplication.

Most of a layer’s work is one calculation: a matrix of inputs multiplied by a matrix of weights.

One cell at a time

Each output cell pairs a row of inputs with a column of weights, multiplies them pair by pair and adds the results. The first cell here takes six multiplications.

Every cell at once

No cell depends on another, so they can all be worked out at the same moment. Hardware for AI is built around that independence.

At real size

A real layer multiplies matrices with thousands of rows and columns. Writing one token takes roughly two operations for every parameter in the model.

For a 70-billion-parameter model, that is about 140 billion operations per token. Kaplan et al., 2020

Chapter 12 of 22

Why graphics chips

A job made of billions of identical sums needs many hands, not a few clever ones.

A CPU has a handful of powerful cores, built to race through complicated tasks one after another.

A few at a time

Given thousands of independent sums, a CPU works through them a few at a time.

Thousands at once

A GPU has thousands of simpler cores that run the same instruction on different numbers at once. Designed to colour the pixels of games, the idea suits neural networks just as well.

Tensor cores

Modern GPUs add tensor cores: units that multiply small matrices in a single step, usually at reduced precision.

An NVIDIA H100 (SXM) has 16,896 CUDA cores and 528 Tensor Cores. NVIDIA

Chapter 13 of 22

The real bottleneck

Arithmetic is fast. Fetching the numbers to do it on is slow.

Around the compute die sit stacks of high-bandwidth memory, HBM, which hold the model’s weights.

Every weight, every token

To write one token, the chip reads every weight from memory. While the data streams in, the arithmetic units mostly wait.

The arithmetic of waiting

A 70-billion-parameter model at 16 bits is 140 GB. At 3.35 TB/s, one full read takes about 42 thousandths of a second, which caps a single conversation near 24 tokens a second, even if it fitted on one chip.

H100 (SXM): 80 GB of HBM3 at 3.35 TB/s. NVIDIA

Share each read

Serving many requests together lets each read of the weights do many people’s work. That is why AI services batch requests.

Part four: scale

One GPU is not enough. Not even close.

Chapter 14 of 22

Why so many GPUs

Large models outgrow one GPU’s memory, and training needs far more still.

A 70-billion-parameter model at 16 bits needs 140 GB for its weights alone. One GPU holds 80 GB. Try the other sizes below the picture.

Split it

So the model is split across GPUs. Each holds a slice of every layer, and they trade results over fast links after each one.

Training needs more

Training also keeps gradients and optimizer state: about 16 bytes per parameter with the common Adam optimizer, before counting the activations.

16 bytes per parameter for mixed-precision Adam. Rajbhandari et al., 2019

Keep in step

To learn from more text at once, copies of the model train on different batches, then combine their gradients so every copy takes the same step.

Chapter 15 of 22

The training bill

Large models take millions of GPU hours. The only way through is many GPUs at once.

Meta reports that training its largest Llama 3.1 model took 30.84 million GPU hours.

Llama 3.1 405B: 30.84M H100 GPU hours. Meta

On one GPU

Done on a single GPU, that would take about 3,500 years.

On sixteen thousand

Meta trained it on up to 16,384 GPUs at once. Divided perfectly, that is about eleven weeks; real training also loses time to failures and waiting. Move the slider to see other counts.

Up to 16,384 H100 GPUs. Meta, 2024

And the power

Each of those GPUs can draw up to 700 watts: together about 11.5 megawatts, before cooling, networking and the rest of the building.

H100 (SXM): up to 700 W. NVIDIA

Chapter 16 of 22

A team of experts

Some of the largest models use only a small part of themselves for each token.

In the transformer, every token passes through every weight of each feed-forward block. A bigger block means more arithmetic for every token.

A router picks

A mixture-of-experts layer splits that block into several smaller ones, the experts, and adds a small router. For each token the router scores the experts and sends the token to the best few, blending their answers. Mixtral uses eight experts in each layer and picks two.

Mixtral: Jiang et al., 2024

Each token, its own pair

Different tokens take different experts; the router learns its choices during training. They are not tidy subjects: in Mixtral, experts did not split by topic, and the patterns found were closer to grammar.

All in memory, a few at work

Every expert must still sit in GPU memory, but each token only does the arithmetic of the ones it visits. A model can hold far more than it uses on any one token.

Mixtral 8x7B: 47B parameters, 13B used for each token. DeepSeek-V3: 671B parameters, 37B used for each token. Mixtral, DeepSeek-V3

Chapter 17 of 22

Inside the data centre

Zoom out, from one chip to the building it lives in.

One GPU: a compute die, with memory stacked beside it.

A server

Eight GPUs share a server, joined by fast links, such as NVIDIA’s NVLink, so that they can work almost as one.

Racks and networks

Servers fill racks, and racks are wired together by a network that carries results and gradients between thousands of GPUs.

Power in, heat out

Almost all the electricity a data centre uses ends up as heat. Cold air or liquid flows in, hot air or liquid flows out, around the clock.

Part five: back to you

Now, back up to your question.

Chapter 18 of 22

Serving your question

When you press Enter, your question joins many others on a GPU somewhere.

Your question arrives as tokens.

Prefill

The model reads the whole prompt in one pass, every position in parallel. This is the pause before the first word appears.

Decode

Then it writes one token per step. Each step reuses what earlier tokens left in the KV cache, so nothing is computed twice, and the cache grows with the conversation.

Many at once

A server interleaves many conversations and adds new ones as others finish, so every read of the weights serves as many people as it can.

Chapter 19 of 22

Smaller numbers

Store each weight in fewer bits, and the model shrinks, at a small cost in accuracy.

Training usually keeps a 32-bit copy of every weight: precise, and large. For a 70-billion-parameter model, 280 GB.

16 bits

Most models run in 16 bits. Here the weight moves by less than a ten-thousandth.

8 bits

At 8 bits, the values a weight can take thin out. This one lands about 0.02 away, and the model halves again.

4 bits

At 4 bits, sixteen levels times a shared scale stand in for every weight. The model fits in 35 GB, on one GPU, at some cost in quality.

Chapter 20 of 22

Beyond words

The same ideas read pictures, and draw them.

Vision models cut an image into patches: small squares of pixels.

Patches as tokens

Each patch becomes a vector, like a word, and a transformer reads them together.

“An image is worth 16×16 words.” Dosovitskiy et al., 2020

Drawing from noise

Image generators work the other way round. Training teaches a model to remove a little noise from pictures, one step at a time.

Step by step

To draw, it starts from pure noise and removes it step after step, guided by your words, until a picture is left. The steps shown are one picture’s noised versions, played backwards.

Diffusion models: Ho, Jain and Abbeel, 2020

Chapter 21 of 22

Models that use tools

Given tools, a model can look things up and act, not only write.

Your question reaches the model.

Calling a tool

When it lacks something, the model can write a request for a tool, such as a search, instead of an answer. The software around it decides whether that call is allowed.

Results join the context

The tool’s result is added to the model’s context, the same way your question was.

An answer, with its source

With the result in view, the model writes its answer, and can point to where it came from.

Chapter 22 of 22

What it is not

A model predicts likely words. Likely is not the same as true.

An answer can read fluently and confidently.

Likely, not true

Each token was chosen because it was likely in context. Nothing in that process checks the claim against the world.

Check a source

A reliable source settles it: sunlight scatters off the gases in the air, and blue light scatters more than red.

Why is the sky blue? NASA Space Place

The last check is yours

Tools and sources help, but judgment stays with the reader. Ask where a claim comes from before you rely on it.

The way back up

That is how AI works.

You asked: Why is the sky blue?

Air molecules scatter blue light more than red.

Tokens, vectors and attention; a walk downhill through billions of weights; arithmetic on thousands of cores, in buildings full of them; and back to one sentence on your screen.

Sources

The pictures are simplified teaching models drawn in your browser, not a live view of any product. Figures and results come from these sources; illustrative values are marked where they appear.

  1. OpenAI: What are tokens and how to count them?
  2. Mikolov et al., 2013: Efficient Estimation of Word Representations in Vector Space
  3. Hendrycks and Gimpel, 2016: Gaussian Error Linear Units (GELUs)
  4. Vaswani et al., 2017: Attention Is All You Need
  5. Rumelhart, Hinton and Williams, 1986: Learning representations by back-propagating errors
  6. Ouyang et al., 2022: Training language models to follow instructions with human feedback
  7. Rafailov et al., 2023: Direct Preference Optimization: Your Language Model is Secretly a Reward Model
  8. Kaplan et al., 2020: Scaling Laws for Neural Language Models
  9. Rajbhandari et al., 2019: ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
  10. Meta, 2024: The Llama 3 Herd of Models
  11. Meta: Llama 3.1 model card
  12. Meta, 2024: Introducing Meta Llama 3
  13. Jiang et al., 2024: Mixtral of Experts
  14. DeepSeek-AI, 2024: DeepSeek-V3 Technical Report
  15. NVIDIA: H100 Tensor Core GPU specifications
  16. NVIDIA: Hopper architecture whitepaper
  17. Dosovitskiy et al., 2020: An Image is Worth 16x16 Words
  18. Ho, Jain and Abbeel, 2020: Denoising Diffusion Probabilistic Models
  19. NASA Space Place: Why Is the Sky Blue?