Learn AI/AI engineering
LESSON 28 / 36Intermediate 35 min with practice

Latency, caching, and bounded work

Control the amount of work before optimizing individual operations.

WHAT YOU WILL LEARN
  • Break latency into stages
  • Use correct cache keys
  • Set explicit work and cost limits

Bound the work first

Total response time includes validation, network delays, retrieval, inference, and rendering. Instrument the stages before optimizing. A fast average can hide poor tail latency, so record percentiles and timeout rates when you have enough requests.

Limit input size, retrieved passages, generation length, tool calls, and retries. Retry only failures likely to be transient, with a capped attempt count and backoff. Unbounded retries can multiply both latency and cost.

Cache only compatible results

A cache key must include the factors that change the result: normalized input, model version, relevant settings, and data version. Private results also need an identity or tenant boundary. A globally shared cache keyed only by question can leak another user’s information.

A cost estimate is a planning tool, not a billing guarantee. Reserve a worst-case budget before paid work begins, enforce provider output limits, reconcile actual usage, and stop when the allowance is exhausted. This lesson uses symbolic work units and makes no paid requests.

PUT THE IDEA INTO CODE

A small experiment you can run.

Changing the corpus version creates a different key. This in-memory cache is only for public deterministic examples and does not implement private multi-user access.

latency-caching-and-budgets.py
from functools import lru_cache
calls = 0
@lru_cache(maxsize=32)
def retrieve(query, corpus_version):
    global calls
    calls += 1
    return f"Result for {query!r} in {corpus_version}"
print(retrieve("python", "v1"))
print(retrieve("python", "v1"))
print(retrieve("python", "v2"))
print("Actual work:", calls)
print("Cache:", retrieve.cache_info())
Copy code

Save the file, open your terminal in that folder, and run python latency-caching-and-budgets.py. Use python3 or py if required by your installation. Setup guide

What to expect

Three calls produce two computations and one cache hit.

YOUR TURN

Design a bounded retrieval request.

  1. Choose a maximum query length and passage count.
  2. Specify one timeout and a maximum retry count.
  3. List every factor your cache key must include.
Compare with a suggested solution

For a public corpus, include query normalization, corpus version, and retrieval settings. For private corpora, add the trusted tenant identity and invalidate on access changes. Keep failed or partial results distinct from successful responses.

CHECK YOUR UNDERSTANDING

One idea to take with you.

Why include a corpus version in a retrieval cache key?

Make it part of your progress.

Finish the practice and answer the knowledge check to mark this lesson complete.

Go deeper with primary documentation

Optional references for further study. This lesson and its examples were written for Artificials.

Python functools module