KV Cache and Batching: Raising LLM Serving Throughput

LLM serving throughput is decided by two memory-side mechanics before any GPU upgrade matters: how the engine stores the key-value (KV) cache and how it refills batches during decoding. Profiling published with the vLLM paper found that in systems that pre-allocate contiguous cache for each request, only 20.4% to 38.2% of KV cache memory actually holds token states; the rest is fragmentation and over-reserved space that no other request can use. Fixing that waste with paged cache management, then keeping the batch full with iteration-level scheduling, is the highest-leverage throughput work available to a serving team.

Why the KV cache caps throughput

Autoregressive decoding is memory-bandwidth-bound: every forward pass reads the model weights once regardless of batch size, so throughput rises almost linearly with the number of sequences packed into memory. The binding constraint is the KV cache. For the 13B-parameter OPT model, the vLLM authors calculate 800 KB of KV cache per token — 2 (key and value) × 5120 (hidden size) × 40 (layers) × 2 bytes of FP16 — which compounds to roughly 1.6 GB for a single 2048-token request. An 80 GB GPU holds weights plus only a few dozen such requests, and traditional engines make it worse: because deep learning frameworks require contiguous tensors, they reserve each request’s maximum sequence length up front. Measuring across serving stacks, the vLLM team found that existing systems waste 60% to 80% of memory to fragmentation and over-reservation — capacity that would otherwise hold more concurrent sequences and therefore more tokens per second. Compute keeps outpacing capacity too: from A100 to H100 the FLOPS more than doubled while the memory ceiling stayed at 80 GB, so the memory side of the equation tightens with every hardware generation.

PagedAttention fixes memory waste

PagedAttention applies operating-system paging to the KV cache: each sequence’s cache is split into fixed-size blocks that need not be contiguous in GPU DRAM, and a block table maps logical to physical blocks, allocated on demand. Internal fragmentation shrinks to the unfilled tail of the last block, external fragmentation disappears because every block has the same size, and finished requests free their blocks immediately for reuse. The vLLM blog reports waste under 4% in practice, versus the majority share lost under contiguous schemes. Blocks also enable copy-on-write sharing, so parallel sampling and beam search reuse prompt KV state instead of duplicating it. The measured effect is large: across LLaMA, GPT and OPT workloads, vLLM lifted serving throughput by 2–4× at the same latency compared with FasterTransformer and Orca, with the gap widening for longer sequences, larger models and more complex decoding algorithms.

Continuous batching keeps GPUs busy

Paged memory determines how many sequences fit; batching policy determines how full the batch stays. Static batching runs a group of requests to completion: early finishers idle their slots while a straggler keeps decoding, and newly arrived requests wait for the whole batch to drain — padding waste and head-of-line blocking in a single design. Orca replaced this with iteration-level scheduling: the scheduler admits and retires requests at every decode step, so finished slots are backfilled immediately and the running batch stays near its memory-limited maximum. On a GPT-3 175B workload this delivered a 36.9× throughput improvement over NVIDIA FasterTransformer at the same latency, and selective batching is what lets prefill and decode phases coexist in one step. Modern engines extend the idea with chunked prefill, which slices long prompts across scheduler steps so prompt compute does not stall ongoing decodes — trading a bounded per-token latency ceiling for a higher average batch size. The heavier the tail of your output-length distribution, the bigger the win over request-level batching.

Sizing your serving stack

Work the capacity math from the memory budget down. Example: a 13B model in FP16 needs about 26 GB for weights; at gpu_memory_utilization 0.90 an 80 GB GPU leaves roughly 46 GB for KV cache, which at 800 KB per token is about 57,000 tokens — around 55 concurrent sequences averaging 1,024 tokens each. That KV budget, not your request queue, is the real concurrency ceiling, so tune the scheduler against it.

KnobControlsStarting point
gpu_memory_utilizationGPU DRAM reserved for weights plus KV cache0.90, leaving headroom for activations
max_num_seqsMaximum concurrent sequencesDerive from the KV token budget, not the queue
max_num_batched_tokensTokens per scheduler step; chunked prefill budget2,048–8,192 depending on latency target
block_sizeKV cache page size16 (engine default)
enable_prefix_cachingReuse of shared prompt prefixesOn for chat and RAG traffic

Throughput tuning checklist

  1. Benchmark with a trace that mirrors production prompt and output lengths; synthetic uniform traffic hides exactly the variance that batching must absorb.
  2. Set gpu_memory_utilization to reserve nearly all free DRAM, then compute the KV token budget and derive max_num_seqs from it.
  3. Raise max_num_batched_tokens until time-to-first-token under load meets your SLO, using chunked prefill to cap inter-token latency.
  4. Enable prefix caching when prompts share system prefixes; it removes duplicated prefill compute and duplicated KV blocks.
  5. Track TTFT p99 and inter-token latency together — throughput gains that break the ITL budget are not gains.
  6. Re-run the cost model before scaling out: cheaper tokens per second also shifts the fine-tuning vs prompt engineering trade-off, and if you move workloads to managed endpoints, serverless cold-start behaviour changes the latency profile your batching tuned away.

Sources