GPU capacity planning for LLM inference is one equation followed by three checks: multiply the memory each replica needs (model weights plus KV cache per concurrent request) by the number of replicas your token-throughput target demands, then verify the result against your latency SLO, your provider quota, and a failover allowance before anyone commits budget. Compute rarely binds first; memory does. An H100 SXM carries 80 GB of HBM with 3.35 TB/s of memory bandwidth, and autoregressive decoding is bound by that bandwidth long before tensor-core peak matters, because every generated token re-reads weights and cache from HBM. Once the arithmetic settles, compare deployment options against GPU total cost of ownership, on-prem versus cloud — but only after sizing, or you will price the wrong fleet.
Start With the Memory Budget
Each replica first needs room for weights: parameters multiplied by bytes per parameter. A 70-billion-parameter model in FP16 needs roughly 140 GB before a single request arrives — an illustrative calculation from this article, not a measured case. That footprint alone spans multiple GPUs and forces a tensor-parallel layout, which in turn changes scheduling, failure domains and per-node utilization. On top of the weights sits the KV cache, which grows with sequence length and concurrency; naive plans fail here in both directions. Reserve too little and requests queue, retry or spill; reserve too much and you pay for idle HBM across every replica, every hour.
Serving engines attack the waste directly. For decoder-only transformers, memory, not compute, is the binding constraint during generation, so the effective lever is how tightly the engine packs concurrent sequences. PagedAttention, introduced in the SOSP 2023 vLLM paper, treats the KV cache like virtual-memory pages instead of one contiguous preallocation per request; the original evaluation reports that it improves the throughput of popular LLMs by 2-4x at the same level of latency compared with earlier systems such as FasterTransformer and Orca. With a paged engine, capacity planning shifts from worst-case preallocation to measuring the batch size your SLO actually sustains.
Convert Throughput Into GPU Count
Once memory says how many concurrent sequences one replica holds, throughput says how many replicas you need. Run your own benchmark with production-shaped prompts and your real context lengths; vendor marketing numbers describe different conditions than your traffic. The core sizing rule is replicas = ceil(target tokens per second ÷ measured tokens per second per replica), then round up again if the p99 latency at full batch exceeds the SLO, because queueing delay grows nonlinearly near saturation. Build the worksheet before quoting any fleet size:
| Input | Where it comes from | Hypothetical example |
|---|---|---|
| Weights footprint | parameters × bytes per parameter | 70B × 2 bytes ≈ 140 GB (FP16) |
| KV cache per sequence | engine metrics at your context length | measured at 4k context, per request |
| Sequences per replica | free HBM ÷ per-sequence cache | derived from the two rows above |
| Per-replica throughput | your benchmark at SLO latency | tokens/s at p99, not at p50 |
| Replica count | ceil(target ÷ per-replica) | plus one spare, see below |
Two refinements matter in practice. First, prefill and decode have different profiles; if you can split them, size the phases separately or long prompts will starve interactive traffic. Second, if you slice GPUs with Multi-Instance GPU on Hopper parts — the NVIDIA H100 specification lists up to seven instances of 10 GB each on H100 SXM — small models can share one card, but each slice gets a fixed memory and compute fraction, so re-measure rather than dividing the monolithic number.
Quotas Are Not Capacity
A sizing that your subscription cannot launch is a slide, not a plan. On Azure, quota is a credit limit, not a capacity guarantee: it caps what you may request, while actual GPU availability remains a separate constraint, per the Azure Machine Learning quota guidance. The same page states that specialized GPU VM families start with a default quota of zero cores, so NC_A100_v4, NDv2 and similar series require an explicit increase before the first node can exist — a lead time item, not a launch-day item. Sizing must also absorb platform overhead: Azure reserves 20% extra compute resources for upgrades, so a deployment requesting 10 instances must have quota for 12, or the deployment itself fails even when traffic fits. Managed endpoints shift the constraint rather than remove it: Amazon Bedrock controls model inference with quotas on token usage, per the Bedrock quotas documentation, with separate per-model allocations for its runtime and mantle endpoints — so a plan that ignores tokens-per-minute will throttle precisely at peak.
Headroom, Spares and Failover
Steady-state sizing is the floor, not the fleet. Plan one replica you are willing to lose: rolling upgrades, driver patches and node reclamation all remove capacity, and on low-priority or spot capacity preemption can arrive mid-job, which is why checkpointing is part of capacity design rather than an afterthought. Decide explicitly how latency degrades during a loss — fewer in-flight batches, admission control at the gateway, or shedding background work — because the alternative is an uncontrolled queue that breaches SLO for everyone at once. Finally, close the loop with money: once the fleet runs, reconcile what the invoice says against what the plan assumed, using FOCUS 1.4 invoice reconciliation so utilization drift is caught by data rather than by a budget alert.
A Working Capacity Checklist
- Compute the weights footprint for the exact quantization you will deploy.
- Measure KV cache per sequence at your real context length using engine metrics.
- Derive sequences per replica from free HBM, then benchmark per-replica tokens per second at SLO latency.
- Apply replicas = ceil(target ÷ per-replica), then add one spare replica for upgrades and failures.
- Check GPU VM family quotas per region and per subscription; treat zero-default families as a lead-time risk.
- Add the platform overhead your provider documents, including upgrade reserves, to the quota request.
- Re-run the benchmark and the invoice reconciliation every month; capacity plans rot silently.