· inference  · 11 min read

LLM serving KV cache: Interview Answer Framework

LLM serving KV cache: Interview Answer Framework. Complete preparation framework with real questions and model answers.

LLM serving KV cache: Interview Answer Framework. Complete preparation framework with real questions and model answers.

LLM Serving KV Cache: Interview Answer Framework

Answer First

The KV cache (key-value cache) stores the attention keys and values computed for every token already processed in a transformer’s forward pass, so that generating each new token only requires computing attention for that one new token against the cached history instead of recomputing attention over the entire sequence from scratch. The correct interview answer names the memory cost formula, explains why it grows linearly with sequence length and batch size, and connects that growth directly to the two hardest serving problems it causes: GPU memory exhaustion at long context lengths and wasted memory from over-allocating for variable-length sequences — the second problem being exactly what PagedAttention was built to solve.

Scope and Assumptions

This page covers KV cache mechanics and its serving-system implications for autoregressive transformer decoding (GPT-style models). It does not cover training-time attention optimizations (FlashAttention is related but solves a different problem — training and prefill compute efficiency, not cross-request memory management) or non-transformer architectures. The interview format assumed: a systems-focused round asking “explain how you’d reduce serving cost for a high-traffic LLM API,” where KV cache understanding is required to give a technically grounded answer rather than a surface-level “just use a bigger GPU” response.

Core Framework: Why the KV Cache Exists and What It Costs

The mechanism. In transformer self-attention, generating token N requires computing attention scores between token N’s query vector and the key/value vectors of all N-1 previous tokens. Without caching, generating a sequence of length L would require recomputing keys and values for tokens 1 through L-1 at every single decoding step, an O(L²) cost. The KV cache stores each token’s key and value vectors the first time they are computed (during the prefill phase, when the input prompt is processed) and reuses them for every subsequent decoding step, reducing per-token generation cost to O(L) — attention against the cache, not recomputation of the cache.

The memory cost formula. For a given model, KV cache memory per token is:

KV cache bytes per token = 2 (K and V) × num_layers × num_kv_heads × head_dim × bytes_per_param

Example — a 7B parameter model, 32 layers, 32 attention heads, head_dim 128, FP16 (2 bytes):
  = 2 × 32 × 32 × 128 × 2 bytes
  = 524,288 bytes per token ≈ 512 KB per token

At a 4,096-token context length, one sequence's KV cache: 512 KB × 4,096 ≈ 2.1 GB

This is the number that makes long-context serving expensive: doubling the context length doubles KV cache memory per sequence, and serving multiple concurrent sequences multiplies that by batch size. A 7B model’s weights fit in roughly 14GB at FP16, but a batch of 16 concurrent sequences at 4K context can require over 30GB of KV cache alone — more memory than the model weights themselves.

Grouped-query attention (GQA) as a direct mitigation. Models designed for efficient serving reduce num_kv_heads below num_query_heads (grouped-query attention shares key/value heads across multiple query heads), which shrinks the KV cache formula’s num_kv_heads term directly. This is why interview answers that only discuss serving infrastructure and skip model architecture miss half the picture — the KV cache cost is partly an architecture decision made before the model ever reaches a serving system.

Worked Example: Diagnosing a Memory-Bound Serving System

Scenario stated by interviewer: “Our LLM API serves a 13B model on a single 80GB GPU. Under load, we can only serve 8 concurrent requests before running out of memory, and each request has a variable context length between 500 and 8,000 tokens. How do you increase throughput?”

Step 1 — Establish the memory budget. Model weights at FP16: roughly 26GB. Remaining GPU memory for KV cache and activation overhead: roughly 50GB (reserving headroom for CUDA overhead and activation memory during prefill). If naive allocation reserves the maximum context length (8,000 tokens) for every request regardless of actual length, and per-token KV cache cost for this model is roughly 800KB (larger than the earlier 7B example due to more layers/heads), the naive per-request reservation is 8,000 × 800KB ≈ 6.4GB — meaning only ~8 requests fit in 50GB even though most requests use far less than the max context. This is the exact bug in the scenario: static, worst-case memory allocation wastes memory on every request shorter than the maximum.

Step 2 — Name the fix: PagedAttention. PagedAttention (introduced in the vLLM serving system) manages the KV cache the way an operating system manages virtual memory: it divides the cache into fixed-size blocks and allocates blocks to a sequence only as that sequence actually generates tokens, rather than reserving the maximum possible length up front. This eliminates the internal fragmentation of static allocation — a request that only uses 500 tokens only consumes 500 tokens’ worth of cache blocks, not 8,000 tokens’ worth.

Step 3 — Quantify the expected improvement. With paged, on-demand allocation instead of static worst-case reservation, the same 50GB budget can serve a batch whose average context length (not max) determines effective capacity. If the request mix in this scenario averages 1,500 tokens instead of the 8,000-token worst case, effective concurrent capacity increases roughly in proportion to the ratio between worst-case and average allocation — in this scenario, several times the original 8 concurrent requests, though the exact multiplier depends on the real request-length distribution and must be measured, not assumed.

Step 4 — Name the secondary lever: continuous batching. Static batching processes a fixed batch together, waiting for the slowest sequence in the batch to finish before starting new requests, wasting GPU cycles on sequences that finished early. Continuous batching (also called in-flight batching) admits new requests into a running batch as soon as any sequence completes, keeping the GPU’s batch dimension full at all times. This is a distinct throughput lever from PagedAttention: PagedAttention solves memory utilization, continuous batching solves compute utilization, and a production-grade serving stack needs both.

Trade-offs Table: KV Cache Management Strategies

StrategyWhat it solvesCostWhen to use
Static max-length allocationSimplicity — trivial to implementWastes memory on any request shorter than the reserved maximumOnly acceptable for low-traffic systems with uniform, short context lengths
PagedAttention (block-based paging)Memory fragmentation from variable-length sequencesImplementation complexity — requires a serving framework that supports it (vLLM, TensorRT-LLM)High-traffic systems with variable context lengths — the default choice for production LLM serving in 2026
KV cache quantization (INT8/FP8 cache)Reduces per-token cache memory further, independent of allocation strategySmall accuracy degradation, must be validated per model/taskCombine with PagedAttention when memory is still the binding constraint after paging alone
Grouped-query attention (GQA) at the model architecture levelReduces num_kv_heads, shrinking cache size at the sourceRequires choosing or training a GQA-native model — cannot be retrofitted onto an existing MHA (multi-head attention) checkpoint without retrainingModel-selection decision, made before serving infrastructure is built
Sliding-window / cache evictionBounds cache growth for very long contexts by discarding old tokens’ KV entriesLoses access to evicted context — unsuitable for tasks needing full-document recallLong-running conversational agents where only recent context matters, not document QA

Decision Rubric

If the serving bottleneck is memory (OOM errors under load, low concurrent request count): diagnose whether allocation is static or paged first — this is usually the highest-leverage single fix, ahead of hardware upgrades.

If the serving bottleneck is throughput with memory headroom available: check whether batching is static or continuous — GPU utilization gaps from waiting on the slowest sequence in a static batch are a common and cheap-to-fix throughput loss.

If context lengths are uniformly short (under 1-2K tokens) and traffic is low: static allocation and static batching may be sufficient, and adopting a more complex serving framework may not be worth the operational overhead — say this explicitly in an interview to show you are not reaching for complexity by default.

If accuracy-sensitive and memory-constrained simultaneously: KV cache quantization must be evaluated per-task, not assumed safe — state that you would run an eval set comparing FP16 vs quantized cache outputs before shipping, not just cite the memory savings.

Interview Scorecard

SignalWeak answerStrong answer
MechanismSays “KV cache speeds things up” with no formulaStates the memory-per-token formula and computes a concrete number for the scenario
Root cause diagnosisJumps to “add more GPUs”Identifies static allocation as the specific failure before proposing a fix
Solution specificityNames “vLLM” with no explanation of the mechanismExplains PagedAttention’s block-based allocation and why it fixes fragmentation
BreadthDiscusses only one leverDistinguishes memory-layer fixes (paging, quantization) from compute-layer fixes (continuous batching)

Book Sample

The 0→1 AI Engineer Interview Playbook (ASIN B0H2CML9XD) includes this KV cache memory formula and the full PagedAttention worked example as a standalone systems-design drill, with follow-up questions an interviewer is likely to ask after the initial diagnosis. The 0→1 Machine Learning Engineer Interview Playbook (ASIN B0H256Z1MF) covers the underlying transformer attention mechanics (query/key/value projections, multi-head attention) for candidates who need that foundation refreshed before tackling serving-system questions.

Get the AI Engineer Interview Playbook: /go/B0H2CML9XD?source=ai-engineers-blog&page=aie-kvcache-001

Get the Machine Learning Engineer Interview Playbook: /go/B0H256Z1MF?source=ai-engineers-blog&page=aie-kvcache-001

Common Follow-Up Questions and How to Handle Them

“Why not just always allocate the maximum context length and accept the waste — isn’t paging added complexity worth avoiding?” Answer with the concrete cost trade-off, not a general appeal to efficiency. Static allocation is defensible only when the request-length distribution is tight (most requests near the same length) — in that case, waste is small and paging’s implementation complexity is not worth adopting. When the distribution is wide (as in the worked example, 500 to 8,000 tokens), static allocation wastes memory proportional to the gap between each request’s actual length and the reserved maximum, and that waste directly caps concurrent throughput. State that the decision should be driven by measuring the actual request-length distribution, not assumed.

“How does KV cache interact with speculative decoding?” Speculative decoding generates several candidate tokens using a smaller draft model, then verifies them in a single forward pass of the full model — this requires maintaining two KV caches simultaneously (draft model and target model) and reconciling them when the draft’s guesses are accepted or rejected, adding memory overhead on top of the target model’s own cache. Naming this interaction shows you understand that serving optimizations compose and sometimes compete for the same memory budget, not that each one is a drop-in independent win.

“What happens to the KV cache when a multi-turn conversation continues across separate API calls?” If the serving system does not persist the KV cache between calls, every new turn re-runs the full prefill over the entire conversation history, which is wasteful and adds latency proportional to conversation length. Production systems address this with prefix caching (also called context caching by some providers) — reusing the KV cache for the shared prefix (system prompt plus prior conversation turns) across requests, computing only the new turn’s keys and values. State the trade-off: prefix caching requires the serving system to keep cache entries alive between requests, which raises memory pressure on the server versus a fully stateless request model, and requires careful cache eviction policy design so long-idle conversations do not permanently occupy GPU memory.

Worked Example: Estimating Concurrent Capacity Before and After a Fix

A senior-level follow-up asks you to put an actual number on the improvement, not just name PagedAttention. Continuing the scenario from the worked example above:

Given: 50GB available for KV cache, per-token cache cost ≈ 800KB, request lengths uniformly
distributed between 500 and 8,000 tokens, average ≈ 4,250 tokens.

Static allocation (reserves max length per request):
  Memory per request = 8,000 tokens × 800KB = 6.4GB
  Concurrent capacity = 50GB / 6.4GB ≈ 7-8 requests (matches the scenario's stated 8)

Paged allocation (reserves actual length per request, using average as an estimate):
  Memory per request ≈ 4,250 tokens × 800KB ≈ 3.4GB
  Concurrent capacity = 50GB / 3.4GB ≈ 14-15 requests

Estimated improvement: roughly 1.8x more concurrent requests from paging alone, before
adding continuous batching's compute-utilization gains on top.

Present this kind of back-of-envelope math unprompted in a systems interview — it is a stronger signal than naming the right tool, because it shows you can reason quantitatively about whether a proposed fix is worth the engineering effort before committing to build it.

Sources and Freshness

KV cache memory formula and PagedAttention mechanism are documented in the vLLM project’s published paper (“Efficient Memory Management for Large Language Model Serving with PagedAttention”) and standard transformer architecture references. Numeric examples in this page are illustrative calculations based on published model architecture parameters, not a specific vendor benchmark. Last reviewed: 2026-07-15. Next scheduled review: quarterly.

If you’re actively preparing for this process, the 0→1 AI Engineer Playbook covers the judgment frameworks, real question patterns, and structured answers this article draws on — useful when you want a complete preparation system rather than scattered tips.

    Share:
    Back to Blog

    Related Posts

    View All Posts »