· inference · 11 min read
LLM serving KV cache: Interview Answer Framework
LLM serving KV cache: Interview Answer Framework. Complete preparation framework with real questions and model answers.
LLM Serving KV Cache: Interview Answer Framework
Answer First
The KV cache (key-value cache) stores the attention keys and values computed for every token already processed in a transformer’s forward pass, so that generating each new token only requires computing attention for that one new token against the cached history instead of recomputing attention over the entire sequence from scratch. The correct interview answer names the memory cost formula, explains why it grows linearly with sequence length and batch size, and connects that growth directly to the two hardest serving problems it causes: GPU memory exhaustion at long context lengths and wasted memory from over-allocating for variable-length sequences — the second problem being exactly what PagedAttention was built to solve.
Scope and Assumptions
This page covers KV cache mechanics and its serving-system implications for autoregressive transformer decoding (GPT-style models). It does not cover training-time attention optimizations (FlashAttention is related but solves a different problem — training and prefill compute efficiency, not cross-request memory management) or non-transformer architectures. The interview format assumed: a systems-focused round asking “explain how you’d reduce serving cost for a high-traffic LLM API,” where KV cache understanding is required to give a technically grounded answer rather than a surface-level “just use a bigger GPU” response.
Core Framework: Why the KV Cache Exists and What It Costs
The mechanism. In transformer self-attention, generating token N requires computing attention scores between token N’s query vector and the key/value vectors of all N-1 previous tokens. Without caching, generating a sequence of length L would require recomputing keys and values for tokens 1 through L-1 at every single decoding step, an O(L²) cost. The KV cache stores each token’s key and value vectors the first time they are computed (during the prefill phase, when the input prompt is processed) and reuses them for every subsequent decoding step, reducing per-token generation cost to O(L) — attention against the cache, not recomputation of the cache.
The memory cost formula. For a given model, KV cache memory per token is:
KV cache bytes per token = 2 (K and V) × num_layers × num_kv_heads × head_dim × bytes_per_param
Example — a 7B parameter model, 32 layers, 32 attention heads, head_dim 128, FP16 (2 bytes):
= 2 × 32 × 32 × 128 × 2 bytes
= 524,288 bytes per token ≈ 512 KB per token
At a 4,096-token context length, one sequence's KV cache: 512 KB × 4,096 ≈ 2.1 GB
This is the number that makes long-context serving expensive: doubling the context length doubles KV cache memory per sequence, and serving multiple concurrent sequences multiplies that by batch size. A 7B model’s weights fit in roughly 14GB at FP16, but a batch of 16 concurrent sequences at 4K context can require over 30GB of KV cache alone — more memory than the model weights themselves.
Grouped-query attention (GQA) as a direct mitigation. Models designed for efficient serving reduce num_kv_heads below num_query_heads (grouped-query attention shares key/value heads across multiple query heads), which shrinks the KV cache formula’s num_kv_heads term directly. This is why interview answers that only discuss serving infrastructure and skip model architecture miss half the picture — the KV cache cost is partly an architecture decision made before the model ever reaches a serving system.
Worked Example: Diagnosing a Memory-Bound Serving System
Scenario stated by interviewer: “Our LLM API serves a 13B model on a single 80GB GPU. Under load, we can only serve 8 concurrent requests before running out of memory, and each request has a variable context length between 500 and 8,000 tokens. How do you increase throughput?”
Step 1 — Establish the memory budget. Model weights at FP16: roughly 26GB. Remaining GPU memory for KV cache and activation overhead: roughly 50GB (reserving headroom for CUDA overhead and activation memory during prefill). If naive allocation reserves the maximum context length (8,000 tokens) for every request regardless of actual length, and per-token KV cache cost for this model is roughly 800KB (larger than the earlier 7B example due to more layers/heads), the naive per-request reservation is 8,000 × 800KB ≈ 6.4GB — meaning only ~8 requests fit in 50GB even though most requests use far less than the max context. This is the exact bug in the scenario: static, worst-case memory allocation wastes memory on every request shorter than the maximum.
Step 2 — Name the fix: PagedAttention. PagedAttention (introduced in the vLLM serving system) manages the KV cache the way an operating system manages virtual memory: it divides the cache into fixed-size blocks and allocates blocks to a sequence only as that sequence actually generates tokens, rather than reserving the maximum possible length up front. This eliminates the internal fragmentation of static allocation — a request that only uses 500 tokens only consumes 500 tokens’ worth of cache blocks, not 8,000 tokens’ worth.
Step 3 — Quantify the expected improvement. With paged, on-demand allocation instead of static worst-case reservation, the same 50GB budget can serve a batch whose average context length (not max) determines effective capacity. If the request mix in this scenario averages 1,500 tokens instead of the 8,000-token worst case, effective concurrent capacity increases roughly in proportion to the ratio between worst-case and average allocation — in this scenario, several times the original 8 concurrent requests, though the exact multiplier depends on the real request-length distribution and must be measured, not assumed.
Step 4 — Name the secondary lever: continuous batching. Static batching processes a fixed batch together, waiting for the slowest sequence in the batch to finish before starting new requests, wasting GPU cycles on sequences that finished early. Continuous batching (also called in-flight batching) admits new requests into a running batch as soon as any sequence completes, keeping the GPU’s batch dimension full at all times. This is a distinct throughput lever from PagedAttention: PagedAttention solves memory utilization, continuous batching solves compute utilization, and a production-grade serving stack needs both.
Trade-offs Table: KV Cache Management Strategies
| Strategy | What it solves | Cost | When to use |
|---|---|---|---|
| Static max-length allocation | Simplicity — trivial to implement | Wastes memory on any request shorter than the reserved maximum | Only acceptable for low-traffic systems with uniform, short context lengths |
| PagedAttention (block-based paging) | Memory fragmentation from variable-length sequences | Implementation complexity — requires a serving framework that supports it (vLLM, TensorRT-LLM) | High-traffic systems with variable context lengths — the default choice for production LLM serving in 2026 |
| KV cache quantization (INT8/FP8 cache) | Reduces per-token cache memory further, independent of allocation strategy | Small accuracy degradation, must be validated per model/task | Combine with PagedAttention when memory is still the binding constraint after paging alone |
| Grouped-query attention (GQA) at the model architecture level | Reduces num_kv_heads, shrinking cache size at the source | Requires choosing or training a GQA-native model — cannot be retrofitted onto an existing MHA (multi-head attention) checkpoint without retraining | Model-selection decision, made before serving infrastructure is built |
| Sliding-window / cache eviction | Bounds cache growth for very long contexts by discarding old tokens’ KV entries | Loses access to evicted context — unsuitable for tasks needing full-document recall | Long-running conversational agents where only recent context matters, not document QA |
Decision Rubric
If the serving bottleneck is memory (OOM errors under load, low concurrent request count): diagnose whether allocation is static or paged first — this is usually the highest-leverage single fix, ahead of hardware upgrades.
If the serving bottleneck is throughput with memory headroom available: check whether batching is static or continuous — GPU utilization gaps from waiting on the slowest sequence in a static batch are a common and cheap-to-fix throughput loss.
If context lengths are uniformly short (under 1-2K tokens) and traffic is low: static allocation and static batching may be sufficient, and adopting a more complex serving framework may not be worth the operational overhead — say this explicitly in an interview to show you are not reaching for complexity by default.
If accuracy-sensitive and memory-constrained simultaneously: KV cache quantization must be evaluated per-task, not assumed safe — state that you would run an eval set comparing FP16 vs quantized cache outputs before shipping, not just cite the memory savings.
Interview Scorecard
| Signal | Weak answer | Strong answer |
|---|---|---|
| Mechanism | Says “KV cache speeds things up” with no formula | States the memory-per-token formula and computes a concrete number for the scenario |
| Root cause diagnosis | Jumps to “add more GPUs” | Identifies static allocation as the specific failure before proposing a fix |
| Solution specificity | Names “vLLM” with no explanation of the mechanism | Explains PagedAttention’s block-based allocation and why it fixes fragmentation |
| Breadth | Discusses only one lever | Distinguishes memory-layer fixes (paging, quantization) from compute-layer fixes (continuous batching) |
Book Sample
The 0→1 AI Engineer Interview Playbook (ASIN B0H2CML9XD) includes this KV cache memory formula and the full PagedAttention worked example as a standalone systems-design drill, with follow-up questions an interviewer is likely to ask after the initial diagnosis. The 0→1 Machine Learning Engineer Interview Playbook (ASIN B0H256Z1MF) covers the underlying transformer attention mechanics (query/key/value projections, multi-head attention) for candidates who need that foundation refreshed before tackling serving-system questions.
Get the AI Engineer Interview Playbook: /go/B0H2CML9XD?source=ai-engineers-blog&page=aie-kvcache-001
Get the Machine Learning Engineer Interview Playbook: /go/B0H256Z1MF?source=ai-engineers-blog&page=aie-kvcache-001
Common Follow-Up Questions and How to Handle Them
“Why not just always allocate the maximum context length and accept the waste — isn’t paging added complexity worth avoiding?” Answer with the concrete cost trade-off, not a general appeal to efficiency. Static allocation is defensible only when the request-length distribution is tight (most requests near the same length) — in that case, waste is small and paging’s implementation complexity is not worth adopting. When the distribution is wide (as in the worked example, 500 to 8,000 tokens), static allocation wastes memory proportional to the gap between each request’s actual length and the reserved maximum, and that waste directly caps concurrent throughput. State that the decision should be driven by measuring the actual request-length distribution, not assumed.
“How does KV cache interact with speculative decoding?” Speculative decoding generates several candidate tokens using a smaller draft model, then verifies them in a single forward pass of the full model — this requires maintaining two KV caches simultaneously (draft model and target model) and reconciling them when the draft’s guesses are accepted or rejected, adding memory overhead on top of the target model’s own cache. Naming this interaction shows you understand that serving optimizations compose and sometimes compete for the same memory budget, not that each one is a drop-in independent win.
“What happens to the KV cache when a multi-turn conversation continues across separate API calls?” If the serving system does not persist the KV cache between calls, every new turn re-runs the full prefill over the entire conversation history, which is wasteful and adds latency proportional to conversation length. Production systems address this with prefix caching (also called context caching by some providers) — reusing the KV cache for the shared prefix (system prompt plus prior conversation turns) across requests, computing only the new turn’s keys and values. State the trade-off: prefix caching requires the serving system to keep cache entries alive between requests, which raises memory pressure on the server versus a fully stateless request model, and requires careful cache eviction policy design so long-idle conversations do not permanently occupy GPU memory.
Worked Example: Estimating Concurrent Capacity Before and After a Fix
A senior-level follow-up asks you to put an actual number on the improvement, not just name PagedAttention. Continuing the scenario from the worked example above:
Given: 50GB available for KV cache, per-token cache cost ≈ 800KB, request lengths uniformly
distributed between 500 and 8,000 tokens, average ≈ 4,250 tokens.
Static allocation (reserves max length per request):
Memory per request = 8,000 tokens × 800KB = 6.4GB
Concurrent capacity = 50GB / 6.4GB ≈ 7-8 requests (matches the scenario's stated 8)
Paged allocation (reserves actual length per request, using average as an estimate):
Memory per request ≈ 4,250 tokens × 800KB ≈ 3.4GB
Concurrent capacity = 50GB / 3.4GB ≈ 14-15 requests
Estimated improvement: roughly 1.8x more concurrent requests from paging alone, before
adding continuous batching's compute-utilization gains on top.
Present this kind of back-of-envelope math unprompted in a systems interview — it is a stronger signal than naming the right tool, because it shows you can reason quantitatively about whether a proposed fix is worth the engineering effort before committing to build it.
Sources and Freshness
KV cache memory formula and PagedAttention mechanism are documented in the vLLM project’s published paper (“Efficient Memory Management for Large Language Model Serving with PagedAttention”) and standard transformer architecture references. Numeric examples in this page are illustrative calculations based on published model architecture parameters, not a specific vendor benchmark. Last reviewed: 2026-07-15. Next scheduled review: quarterly.
Recommended Resource
If you’re actively preparing for this process, the 0→1 AI Engineer Playbook covers the judgment frameworks, real question patterns, and structured answers this article draws on — useful when you want a complete preparation system rather than scattered tips.