Inside a streaming answer: prefill, decode, and the growing cache
Follow tokens through an animated inference pipeline.

A chat interface makes an AI answer look like a sentence gradually appearing on a screen. Behind that familiar effect are several different kinds of work. The server receives and prepares an input, the model processes the available context, and a generation loop produces a continuation. The interface then turns arriving data into visible text. Those stages need not share the same bottleneck.
The animated figure in this guide follows one simplified request. Scroll through it to move forward, scroll back to reverse the explanation, or use the controls to inspect a stage directly. The moving shapes explain information flow. They are not a timing trace from a commercial model, and their speed does not represent tokens per second.
The second figure isolates one resource that often surprises developers: the key-value cache. Change the context length and number of simultaneous sequences to see an illustrative memory estimate. The architecture assumptions are shown beside the result so the number can be understood rather than mistaken for a measurement of a named service.
Follow one request through the system
One request, from prompt to answer
Reduced motion is on. Choose a stage or move the slider to explore.
Assemble the request
Instructions, the conversation and the question become the available prompt. Blocks are schematic, not actual tokenizer output.
Read every stage without animation
- Assemble the request
Instructions, the conversation and the question become the available prompt. Blocks are schematic, not actual tokenizer output.
- Process the available context
The model processes prompt positions under its attention mask. This is different from selecting future output tokens.
- Retain reusable attention state
Keys and values for earlier positions can be reused. The cache is neither verified knowledge nor permanent user memory.
- Choose the next continuation
A selected token becomes context for the next iteration. Each new position can add cache state.
- Stream, validate, finish
The interface receives output. Completion, cancellation and validation remain application responsibilities.
Begin with the input. An application may assemble system instructions, earlier conversation, retrieved material, and the latest user message. A tokenizer converts text into the model's token representation. A visible word can map to multiple tokens, and a token can contain only part of a word or include punctuation and whitespace. The blocks in the figure are schematic units, not an actual tokenizer's output.
The request then enters model computation. In a conventional autoregressive transformer workflow, the available prompt is processed before the system repeatedly chooses continuations. Earlier positions supply information that later positions can use under the model's attention rules. The figure separates prompt processing from the repeated generation loop because their operational behavior differs.
Prefill is work on the available prompt
Prefill processes the input sequence and prepares representations used when generating a continuation. Many positions in the already available prompt can be processed together, subject to the architecture's causal mask and implementation. This differs from waiting for a future generated token whose identity has not yet been chosen.
A longer prompt can require more computation and memory traffic, but its effect on elapsed time depends on implementation, hardware, batching, and context management. The animation therefore does not make a long prompt take a proportionally exact number of seconds. A conceptual flow should not smuggle in a performance claim that only a real measurement could support.
For an application, distinguish time waiting in a queue from time spent processing the prompt. A busy service can make an otherwise ordinary request feel slow before substantial model work starts. Network connection setup, routing, and application preparation can also contribute to the delay before the first visible output.
Decode depends on what was just produced
During autoregressive decoding, the model produces a distribution for a continuation and the generation procedure selects a token. That token becomes part of the context for the next step. The loop continues until a stopping condition, output limit, or application boundary is reached. The resulting text arrives incrementally only if the server and client expose it that way.
The distinction between tokens and readable text matters here. A network event does not always contain exactly one token, and a token does not always produce a complete visible word. Some interfaces buffer partial output for display. Counting animation frames or text chunks in the browser is therefore not automatically the same as measuring the model's decoding rate.
An application should also know what ends the loop. A natural stopping token, a token budget, a timeout, and a tool call are different events. If the system treats every interrupted answer as complete, a smooth stream can conceal missing reasoning, unfinished instructions, or invalid structured output.
Reuse earlier attention state
The key-value cache stores attention-related representations for earlier positions so an implementation can reuse them during later decoding steps. The useful intuition is avoiding repeated preparation of the same historical information. It is not a database of verified facts, a permanent memory of the user, or a guarantee that the model interpreted those positions correctly.
Hugging Face's cache documentation describes several implementation strategies, including dynamic and static caches and approaches involving offloading or quantization. Those strategies have different tradeoffs. The existence of a cache does not imply that every runtime allocates memory in the same way or supports the same optimizations.
Hugging Face: Technical documentation
The diagram represents cached positions as stored cells. As decoding adds positions, the represented history grows. This is intentionally separate from the model weights, which are another major part of memory use. Confusing weights and cache can make a deployment estimate appear plausible while omitting the resource that grows with active conversations.
Turn the memory assumptions into a calculation
What makes the cache grow?
Each tile represents 0.5 GiB, rounded up for display; the numeric result is exact for these assumptions.
2 × 32 × 8 × 128 × 2 × cached tokens × sequences ÷ 1,073,741,824
The calculator uses a hypothetical transformer with thirty-two layers, eight key-value heads per layer, a head dimension of one hundred twenty-eight, and two bytes per stored element. It accounts for both keys and values. Multiplying those factors gives one hundred thirty-one thousand seventy-two bytes per cached token per sequence before allocator and other overhead.
At eight thousand one hundred ninety-two cached tokens, that illustrative raw cache occupies one gibibyte for one sequence. Two such sequences require two gibibytes under the same assumptions. These are arithmetic results for the displayed hypothetical architecture, not memory figures for any named model in the catalog. Different head counts, precision, attention windows, or cache-sharing behavior change the answer.
Move the sliders and watch the cache tiles and total change together. The relationship being taught is the dependence on stored positions and simultaneous sequences under fixed architecture assumptions. It does not imply that total GPU memory, total latency, or billed cost follows the same simple proportional rule.
Allocation is different from useful content
Even a correct count of raw key-value elements is not a complete serving-memory plan. Implementations may reserve capacity, align blocks, carry metadata, or retain temporary workspaces. Model weights and activations consume additional memory. The difference between allocated capacity and currently useful content can become important when requests have varied lengths.
The PagedAttention research studies memory management for language-model serving and presents an approach inspired by paging. Its central relevance here is the organization of changing cache requirements across requests. It is evidence that allocation strategy matters, not a universal promise that any application will obtain the paper's reported performance under arbitrary conditions.
Read the original research paper on arXiv
When testing a server, include short and long requests together. A benchmark containing only equal-length prompts may hide the fragmentation, scheduling, and cancellation behavior that appears in actual traffic. Measure the workload you expect to serve rather than only the shape that gives the cleanest throughput number.
Streaming improves feedback, not correctness
Showing early output can reassure a user that work is progressing. It can also let them stop an irrelevant answer before it consumes the full output budget. These are interaction benefits. They do not establish that the eventual answer is accurate, grounded, or complete.
For some tasks, displaying unverified partial output is undesirable. An application producing a structured record may need to wait until the object is complete and validated. A tool-using assistant may need to distinguish provisional narration from the result of a completed action. The right streaming policy follows the task's meaning, not just the availability of a streaming API.
Design the interface to communicate interruption. If the connection ends early, preserve enough state to say that the response did not finish. Avoid turning the final visible fragment into a confident answer simply because it arrived without a browser exception. Completion is an application condition as well as a transport event.
Measure the phases separately
For a useful latency investigation, record when the application submits the request, when the first useful output arrives, and when the task finishes. Where the serving stack exposes them, separate queueing, prefill, and decoding measurements. Keep the definitions consistent across runs and providers.
Use several prompt lengths and output lengths. A request with a large document and a short answer stresses a different balance from a short instruction asking for a long draft. Report distributions and difficult cases, not only an average that conceals a tail of slow requests. Include failures and timeouts in the outcome accounting.
If a change appears to improve performance, verify that it did not silently shorten answers, remove required context, or change the accepted-task rate. Faster completion of an easier or incomplete task is not necessarily an improvement in the original workflow. The benchmark must preserve what success means.
Give cancellation a real destination
A stop button should do more than hide incoming text. The application needs to propagate cancellation through its own work and, where supported, to the upstream request. It should then release local resources and stop scheduling dependent steps. Whether already performed provider work remains billable depends on the service's actual terms.
Cancellation also interacts with retries. Automatically starting a second request after a client disconnect can turn a user trying to stop work into more work. Give each request a lifecycle and record whether it completed, failed, or was cancelled. That makes resource behavior easier to understand when the system becomes busy.
The animation ends at an answer, but a production request ends at an explicit state. Keeping those states clear helps both the interface and the infrastructure. It is one of the simplest ways to prevent a visually polished streaming experience from hiding operational confusion.
Read the moving picture as a map
Return to the request diagram and pause at each stage. Ask what information exists, what work can happen together, what depends on the previous output, and what state must remain available. Those questions turn a moving illustration into a useful model of the system.
The picture deliberately leaves out many hardware details. Its purpose is to separate concepts that a chat box merges together: preparation, prompt processing, reuse, generation, and delivery. Once those pieces are distinct, memory estimates, latency measurements, and interface decisions become much easier to discuss precisely.