Back to News & insightsEngineering

The KV cache: why long conversations consume serving memory

Understand the difference between model weights and attention state, then design a capacity test that includes long prompts and concurrent users.

Editorial guide · Updated September 27, 2026 · 4 min read
Translucent memory tiles follow a curved rail beside a dark processor.

A language model can fit on a machine and still run out of memory when people start using it. Loading weights is only the beginning. Active requests need working space, and long conversations can make that space a substantial part of the deployment.

For many autoregressive transformer systems, a key-value cache is part of this working state. Understanding its role helps explain why a successful local demonstration can behave very differently when several users arrive together.

What the cache keeps

During generation, attention uses representations of earlier tokens. A KV cache stores relevant key and value tensors so the system can reuse them instead of recomputing all previous states for every new token. It is internal numerical state, not a saved final answer or a searchable record of facts.

The PagedAttention research addresses the memory-management problem this creates in serving systems. It describes allocation and sharing techniques intended to reduce wasted cache memory. Its published experiments belong to specified workloads and implementations; they do not establish a universal speed multiplier for every deployment.

Why one long request changes capacity

Consider an illustrative study assistant. Most students ask a paragraph-length question, while a few upload an entire course pack. If capacity planning uses only the average prompt, the course-pack requests may consume far more working memory than the plan anticipates.

Memory demand depends on the architecture, cache representation, sequence lengths, and active requests. Different attention designs can change the relationship. Use measurements from the actual serving stack rather than treating one online formula as valid for every model.

Output length matters too. A request that starts small can grow as generation continues. Reserve room for the allowed completion, and account for other workloads sharing the device. An input limit without an output policy leaves the operating envelope only partly defined.

Distinguish three kinds of reuse

An attention cache within one request, shared-prefix reuse across requests, and an application cache of finished answers solve different problems. They may coexist, but enabling one does not imply that the others exist.

The study assistant might reuse computation for a common instruction prefix. That does not authorize sharing one student's private conversation with another. Likewise, finding a previous finished answer in an application cache says nothing about how the underlying model manages attention memory.

Write separate policies for each layer. Identify what can be reused, how identity is checked, and what makes an entry stale. Measure their effects separately so an apparent performance gain is not attributed to the wrong mechanism.

Build a realistic capacity experiment

Start with a fixed model revision and serving configuration. Prepare short, medium, and long prompts drawn from the expected product shape. Then vary concurrent requests while keeping the mix visible. A run composed entirely of identical prompts can overstate benefits available to real users.

For each run, record arrival rate, queue delay, time to first useful output, completion time, peak memory, and errors. Include cancellations. A student who leaves the page should not leave an abandoned job consuming scarce capacity indefinitely.

Repeat with a cold service and with a warmed service. Note whether requests reuse prefixes. The goal is not to manufacture the highest throughput number; it is to identify a safe range in which the application remains predictable.

Make overload behavior deliberate

Suppose the experiment shows that several simultaneous long documents make short questions wait unacceptably. Possible responses include separate queues, smaller document allowances, retrieval before generation, or a lower concurrency limit for expensive jobs.

Each choice has a user-facing consequence. Retrieval can omit evidence; queues add waiting; tighter limits exclude some tasks. Test the resulting answers and explain the relevant limit before users submit a document. Do not silently truncate input and return an answer that appears comprehensive.

If capacity is exhausted, reject or queue work clearly. Automatic retries without a budget can worsen overload. Preserve the student's question and show whether any accepted job is still running before offering another submission.

A capacity number needs a workload definition

Report capacity as a tested workload envelope, not simply a number of users. Ten people asking short questions and ten people analyzing long documents are different demands. Keep the assumptions beside any throughput result.

Revisit that envelope when changing context limits, model architecture, precision, or scheduling. Weight size remains useful, but production capacity depends on everything that must remain alive while an answer is being produced.

Research background

Read the original research paper on arXiv

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.