· inference · 12 min read
LLM serving continuous batching: Interview Answer Framework
LLM serving continuous batching: Interview Answer Framework. Complete preparation framework with real questions and model answers.
LLM Serving Continuous Batching: Interview Answer Framework
Answer First
Continuous batching answers the interview question “how do you serve an LLM efficiently at scale” by replacing static, request-level batching with iteration-level scheduling: the serving engine adds and removes sequences from a running batch at every decoding step instead of waiting for a fixed batch to fully finish. This keeps GPU utilization high even when requests have wildly different output lengths. A strong answer names the mechanism (per-token scheduling, not per-request), the enabling data structure (paged KV cache), and the two failure modes interviewers probe: head-of-line blocking under static batching, and memory fragmentation under naive dynamic batching.
Scope and Assumptions
This page is for a mid-to-senior AI engineer interview question of the form: “Walk me through how you would serve a 7B-70B parameter LLM for a production chat product with p99 latency requirements.” It assumes the candidate already knows what autoregressive decoding is (one token generated per forward pass, conditioned on all prior tokens) and what a KV cache is (cached key/value tensors from prior tokens so each new token does not require recomputing attention over the full sequence). It does not cover training-time batching, which is a distinct topic with distinct constraints (fixed sequence length, no incremental token generation, gradient accumulation instead of request scheduling). The examples below use a single-GPU-class deployment (one A100 80GB) to keep the numbers concrete; multi-GPU tensor and pipeline parallelism is referenced but not derived.
Why Static Batching Fails: The Core Framework
Static batching groups N requests, runs the forward pass for all N in lockstep, and returns results only when every sequence in the batch has finished generating (hit EOS or max tokens). The GPU stays busy during the batch, but the batch cannot start a new request until the slowest sequence in the current batch finishes.
Concretely: if you batch 8 requests where 7 produce 20 tokens and 1 produces 500 tokens, the 7 short sequences sit idle on the GPU (their slots are held, wasting memory and doing no useful compute) for 480 extra decoding steps while the engine waits on the one long sequence. This is head-of-line blocking. Under bursty traffic with variable output length (chat, agent tool-calling, chain-of-thought reasoning), this can cut effective throughput by 3-5x compared to the GPU’s raw FLOPS ceiling, because most of the batch is padding-equivalent dead time.
Continuous batching (also called in-flight batching, iteration-level batching, popularized by Orca and productionized in vLLM and TensorRT-LLM) fixes this by scheduling at the token level:
- At every decoding step, the scheduler checks which sequences in the running batch have finished.
- Finished sequences are evicted immediately, freeing their KV cache memory slots.
- Waiting requests in the queue are admitted into the now-open slots, up to a configured max batch size or KV cache budget.
- The next forward pass runs on this reshuffled batch composition.
This means the batch composition changes every single decoding step, not once per full generation. A short request that finishes at token 20 frees its slot for a new request immediately, rather than after the batch’s longest sequence finishes.
The Second Framework Piece: Paged KV Cache
Continuous batching alone is not sufficient if KV cache memory is allocated contiguously per request, because you cannot predict output length in advance. If you pre-allocate for max sequence length (say 4096 tokens) per slot, you waste memory on every request that finishes early — this is internal fragmentation, and it caps your usable batch size long before you exhaust actual token-level memory needs.
PagedAttention (from vLLM) solves this by managing the KV cache like an OS manages virtual memory: KV cache is divided into fixed-size blocks (typically 16 tokens per block), and each sequence’s cache is a list of block pointers rather than one contiguous allocation. New blocks are allocated on demand as a sequence grows, and freed immediately when it finishes. This is the concrete mechanism that lets continuous batching actually scale batch size, because memory is no longer reserved against a worst-case length.
A strong interview answer states this dependency explicitly: “continuous batching needs paged KV cache to avoid trading head-of-line blocking for memory fragmentation” — this is the sentence that separates a candidate who has memorized the term from one who understands the system.
Worked Example: Serving Architecture Walkthrough
Assume the interviewer says: “Design serving for a customer support chatbot, 7B model, p99 latency target 2 seconds for a 200-token response, expected 50 concurrent users at peak.”
A structured answer:
Step 1 — Estimate memory budget. A 7B model in FP16 takes roughly 14GB of weights. On an 80GB A100, that leaves about 66GB for KV cache and activation memory (reserve ~5GB for activations and overhead, leaving ~60GB for KV cache).
Step 2 — Compute KV cache cost per token. For a 7B model (32 layers, 32 heads, head dim 128), KV cache per token is approximately: 2 (K and V) x 32 layers x 32 heads x 128 head_dim x 2 bytes (FP16) = 524,288 bytes ≈ 0.5MB per token.
Step 3 — Convert to max concurrent tokens. 60GB / 0.5MB per token ≈ 120,000 tokens of KV cache capacity available concurrently across all in-flight sequences.
Step 4 — Translate to batch size. If average context + generation is 1,000 tokens per active sequence, that supports roughly 120 concurrent sequences purely on memory — comfortably above the 50-user peak, leaving headroom for burst traffic and longer support conversations with retrieved context.
Step 5 — Choose the serving engine. State a concrete choice: vLLM for its mature PagedAttention implementation and continuous batching scheduler, or TensorRT-LLM’s in-flight batching if the team already standardizes on NVIDIA’s inference stack and needs the extra kernel-level fusion for lower per-token latency.
Step 6 — Add the queueing detail interviewers listen for. Set a max batch size cap independent of memory (e.g., 64 concurrent sequences) to bound per-step compute time, because continuous batching removes the memory-driven ceiling but a very large batch still increases per-step latency for every sequence in it, since the forward pass computes over the full batch. This is a throughput-vs-latency knob, not a bug.
Trade-offs Table
| Dimension | Static Batching | Continuous Batching (no paging) | Continuous Batching + Paged KV Cache |
|---|---|---|---|
| GPU utilization under variable output length | Low (head-of-line blocking) | Medium (memory fragmentation caps batch size) | High |
| Implementation complexity | Low | Medium | High (custom memory manager, block allocator) |
| p99 latency for short requests behind long ones | Poor | Good | Good |
| Memory efficiency | Poor (worst-case pre-allocation) | Poor (still pre-allocates per slot) | Good (block-level allocation) |
| Best fit | Offline batch scoring, fixed-length embeddings | Low-traffic prototypes | Production chat, agents, RAG with variable context length |
| Maturity of open tooling | Any framework | Rare in the wild alone | vLLM, TensorRT-LLM, TGI |
Decision Rubric
Use this rubric to decide what to say when the interviewer changes the constraints:
- If the workload is offline (batch scoring a fixed dataset overnight, no latency SLA): static batching is defensible and simpler to reason about. Say this explicitly rather than defaulting to continuous batching everywhere — interviewers probe for over-engineering as much as under-engineering.
- If output length varies by more than roughly 3x across requests and there is a latency SLA: continuous batching is required, not optional. State the head-of-line blocking math to justify it.
- If concurrent user count is under ~10 and context length is short (under 512 tokens): the memory pressure that paged KV cache solves may not bind yet. Mention it as the next optimization, not a day-one requirement.
- If context length varies significantly (RAG with retrieved documents of variable size, or agent traces that grow with tool calls): paged KV cache is required alongside continuous batching, because contiguous allocation fragmentation gets worse as variance in sequence length increases.
- If asked “what would you change under 10x more traffic”: name multi-GPU tensor parallelism to serve a larger model, then request-level load balancing across replicas, then speculative decoding to cut per-token latency — in that order of marginal impact.
Follow-up Questions to Expect
- “What happens to a sequence’s KV cache blocks if the batch is full and a new higher-priority request arrives?” — answer with preemption strategies: swap to CPU memory or recompute from scratch, and state the latency cost of each.
- “How does continuous batching interact with speculative decoding?” — speculative decoding generates multiple candidate tokens per step using a small draft model, which complicates the per-step batch accounting used by continuous batching schedulers; this is an active systems research area, not a solved default.
- “What’s the failure mode if you set the max batch size too high?” — per-step latency degrades for every sequence in the batch, because the forward pass is a single matrix operation sized to the batch; note this hits p99 latency even though throughput (tokens/sec across the fleet) goes up.
(hypothetical interview scenario — this is a framework for structuring an answer, not a transcript of an actual interview loop.)
Interview Answer Scorecard
| Signal | Weak Answer | Strong Answer |
|---|---|---|
| Names the mechanism | ”It batches requests together" | "It schedules at the token/iteration level, evicting and admitting sequences every decoding step” |
| Explains the failure it fixes | Vague “it’s faster” | Names head-of-line blocking with a concrete example |
| Connects to memory management | Not mentioned | States the paged KV cache dependency and why fragmentation would otherwise cap batch size |
| Does capacity math | Skips to a tool name | Walks through KV cache bytes-per-token to concurrent sequence count |
| States a trade-off | None | Notes that a large max batch size raises throughput but hurts per-step latency |
Common Mistakes Candidates Make on This Question
Mistake 1: Conflating batching with parallelism. Candidates frequently describe continuous batching as “running requests in parallel,” which is imprecise. The forward pass for a batch is a single sequential operation on the GPU — all sequences in the batch move through the same layers together in one pass. What varies is which sequences are included in that single pass at each step, not whether they execute in parallel with each other in a stronger sense. Precision on this point signals genuine understanding versus memorized vocabulary.
Mistake 2: Treating “larger batch size” as strictly better. A candidate who says “just make the batch as large as possible” without acknowledging the latency cost misses half the trade-off. Larger batches raise aggregate throughput (tokens/sec across the fleet) but raise per-step latency for every sequence sharing that batch, since the matrix multiply is sized to the full batch dimension. State both directions.
Mistake 3: Ignoring prefill versus decode asymmetry. The prefill phase (processing the input prompt, computing all its KV cache entries in one pass) is compute-bound and benefits from large batches; the decode phase (generating one token at a time) is memory-bandwidth-bound, since each step only computes one new token per sequence but must read the full KV cache from memory. Continuous batching schedulers that separate or prioritize these two phases differently (chunked prefill, as implemented in more recent vLLM versions) avoid a long prefill request stalling decode steps for other in-flight requests. Mentioning this distinction moves an answer from “textbook correct” to “systems-aware.”
Mistake 4: Skipping the request admission policy. Continuous batching still needs a policy for which waiting request gets the next open slot — first-come-first-served is the simplest default, but production systems often add priority tiers (e.g., paid tier users get preferential admission) or fairness constraints to avoid starving long-running requests indefinitely if the scheduler always prefers short jobs to maximize throughput. A candidate who is asked to design for a multi-tenant system and does not mention admission policy has left a real gap in the design.
How This Extends to Multi-GPU Serving
Everything above assumes a single GPU holds the full model. Once traffic exceeds a single GPU’s capacity, two orthogonal scaling paths open up, and interviewers often push on which to choose first. Tensor parallelism splits each layer’s weight matrices across multiple GPUs, so a single forward pass for one batch is computed jointly across GPUs — this is required when the model itself does not fit on one GPU (a 70B model in FP16 needs roughly 140GB, exceeding a single 80GB A100), and it also reduces per-token latency by parallelizing compute. Data parallelism (replica-based scaling) instead runs full independent copies of the model on separate GPUs or GPU groups, each running its own continuous batching scheduler, with a load balancer routing requests across replicas. The right answer under a “10x more traffic” follow-up is: if the model already fits comfortably on one GPU and the bottleneck is request volume, scale out with replicas behind a load balancer first, since it is operationally simpler and scales near-linearly; only add tensor parallelism when a single GPU cannot hold the model or when per-token latency itself (not just throughput) is the binding constraint.
Relevant Book Sample
The 0→1 AI Engineer Interview Playbook (ASIN B0H2CML9XD) includes a dedicated system design chapter that walks through this exact serving capacity calculation for three model sizes, plus a rubric for how interviewers at infrastructure-heavy teams grade the head-of-line blocking explanation specifically. If your gap is closer to core ML fundamentals (attention mechanics, why KV cache exists at all), The 0→1 Machine Learning Engineer Interview Playbook (ASIN B0H256Z1MF) builds the prerequisite understanding of the transformer forward pass this framework assumes.
Read the relevant chapter: /go/B0H2CML9XD?source=ai-engineers-blog&page=aie-llm-cb-001
For ML fundamentals underlying this answer: /go/B0H256Z1MF?source=ai-engineers-blog&page=aie-llm-cb-001
Sources and Freshness
Technical claims here are grounded in the publicly documented designs of vLLM (PagedAttention, Kwon et al., 2023) and Orca (continuous batching, Yu et al., OSDI 2022), plus NVIDIA’s public TensorRT-LLM in-flight batching documentation. KV cache memory formulas follow the standard transformer decoder parameter accounting (layers x heads x head_dim x 2 for K/V x bytes per element). No proprietary benchmark numbers or company-specific interview data are used. Last verified: 2026-07. Next review: quarterly, or immediately if vLLM or TensorRT-LLM publish a major scheduler architecture change.
Recommended Resource
If you’re actively preparing for this process, the 0→1 AI Engineer Playbook covers the judgment frameworks, real question patterns, and structured answers this article draws on — useful when you want a complete preparation system rather than scattered tips.