GPU capacity planning for LLM inference is one equation followed by three checks: multiply the memory each replica needs (model weights plus KV cache per concurrent request) by the number of replicas your token-throughput target demands, then verify the result against your latency SLO, your provider quota, and a failover …