Serverless inference cold starts are two stacked problems, not one. The platform provisions compute, attaches storage and pulls your container image; only then does the inference engine load model weights into GPU memory, compile execution graphs and accept traffic. On GPU platforms the second stage dominates: as RunPod documents from its own serverless platform, the infrastructure cold start only covers provisioning compute, attaching storage and pulling the container image, while everything after that — loading weights into GPU memory, running compilation and capturing CUDA graphs — is a second, engine-level cold start that you, not the platform, must engineer away. For a 30 GB model, engine initialization alone can exceed five minutes, and every user who hits scale-from-zero pays that wait.
Anatomy of a Model Cold Start
Traditional cold-start advice targets runtime boot: smaller bundles, fewer dependencies, connection reuse. AI workloads break those assumptions because the unit of initialization is gigabytes of weights moving into a scarce, physically attached accelerator. Google’s Cloud Run engineering guide frames the problem as four sequential phases: infrastructure provisioning with GPU allocation and driver injection (~5s), block-level container image streaming (1-2s), engine warm-up, and finally the transfer of weights into VRAM where GPU memory, not CPU, is the binding constraint. A model whose weights do not fit in VRAM silently degrades as tensors spill to system RAM, so sizing the accelerator is a cold-start decision, not just a throughput decision.
Where the Seconds Actually Go
Profile before tuning. The table below maps the phases to their owners and typical magnitudes as documented by the platform vendors:
| Phase | What happens | Typical time | Who controls it |
|---|---|---|---|
| Infrastructure provisioning | GPU allocation, NVIDIA driver injection | ~5s | Platform |
| Image streaming | Only needed container blocks are pulled | 1-2s | Platform + image design |
| Engine initialization | vLLM or Ollama warm-up, CPU-heavy | 5-15s | You |
| Weight load to VRAM | Checkpoint transfer into GPU memory | Model-size dependent | You |
Per Google’s measurements, engine initialization alone takes 5-15s during a Cloud Run AI cold start, before model weights move into VRAM — and that assumes you are not throttled, which is common because teams under-provision vCPU for a phase that is compute-bound, not GPU-bound. Cloud Run’s CPU boost option temporarily doubles vCPU during startup precisely for this phase. Budget your startup probe accordingly: use a high failure threshold (for example 60 checks at a 5-second period) so the platform does not kill a container that is legitimately loading a large checkpoint.
vLLM Cold Start Tuning
The most instructive published experiment is RunPod’s controlled A/B test on its serverless GPU platform: four configuration changes took a 32B FP8 model on two H200s from a 324-second cold start to 91 seconds, a 3.5× improvement with no custom code and no change to request-time latency. The lesson is that compile and graph capture dominate, not raw weight size. Apply the same four levers on any vLLM deployment:
- Persist the compile cache. Point VLLM_CACHE_ROOT at a persistent volume so torch.compile artifacts and CUDA graphs survive worker death; the cache key covers model, GPU architecture and serving flags.
- Prefetch weights. Set the safetensors load strategy to prefetch so checkpoint files enter the OS page cache ahead of materialization instead of lazy memory-mapping.
- Trim CUDA graph capture sizes. Compile only a power-of-two batch ladder (1, 2, 4, 8, 16, 32, 64) up to your real MAX_NUM_SEQS ceiling; vLLM logs one capture line per size so you can see the boot cost directly.
- Consider eager mode. enforce_eager=True trades some steady-state throughput for sub-minute cold starts — the right trade for spiky, low-utilization endpoints.
Two caveats from the same experiment: attaching a network volume pins the endpoint to one datacenter, trading compile time for placement latency, and setting MAX_NUM_SEQS below your true concurrency ceiling changes serving behavior, not just boot behavior. Snapshot-restore features (RunPod FlashBoot) only pay off when the next request lands on the same host with the same image — bursty low-volume endpoints frequently miss and pay the full boot anyway.
Snapshots and Warm Pools
The other mitigation family keeps initialization out of the request path entirely. AWS Lambda SnapStart attacks the classic FaaS cold start: when a version is published, Lambda resumes new execution environments from the cached snapshot instead of initializing them from scratch, restoring an encrypted Firecracker microVM image of already-initialized memory and disk. SnapStart covers Java 11+, Python 3.12+ and.NET 8+ runtimes and helps code-level initialization — dependency loading, framework boot — but it does not eliminate the GPU weight-load phase, so treat it as one layer, not a cure. The blunter instrument is warmth: provisioned concurrency on Lambda, minimum instances on Cloud Run, always-on workers on GPU platforms. Warmth converts a variable cold start into a fixed monthly bill, and vendor guidance converges on the same break-even logic — RunPod’s numbers put the crossover for always-on workers at roughly 25% monthly utilization, below which scale-to-zero plus a fast cold start is cheaper. Cloud Run keeps idle instances alive for about 15 minutes after the last request, so predictable traffic every 10-12 minutes may need no warm pool at all.
A Decision Checklist
Work the problem in this order:
- Measure the cold-start share of latency (INIT duration on Lambda, delayTime on serverless GPU) before spending money; production LLM observability should expose cold versus warm invocations as a first-class metric.
- Persist compile caches and trim CUDA graph capture sizes — the cheapest wins, often 3× on boot time.
- Right-size startup compute: more vCPU (or platform CPU boost) during initialization, tuned startup probes that do not false-positive on port-open.
- Set platform concurrency to the engine’s real parallelism: (model instances × parallel queries) + (model instances × ideal batch size), so one warm instance absorbs bursts instead of triggering scale-out.
- Decide warmth versus scale-to-zero on measured utilization, pricing the warm floor as a fixed cost against the tail-latency benefit, using the same arithmetic as your broader cloud AI cost optimization plan.
- Mask the residual: fire a lightweight health check from the client when a user opens the chat surface, so infrastructure and image phases complete while they are still typing.
Cold starts are an engineering budget, not a platform defect. Platforms own roughly five to seven seconds of the boot; everything beyond that is your engine configuration, your checkpoint layout and your warmth policy. Teams that persist compile artifacts, cap graph capture and measure utilization before buying warm capacity routinely turn multi-minute starts into a bounded worst case their users never feel.