Cloud GPU capacity usually beats owned hardware on total cost of ownership below roughly half of sustained utilization, and loses above it — provided you price power, cooling, staff and hardware refresh honestly. On the cloud side, AWS lists an effective rate of $41.528 per instance-hour for a p5.48xlarge EC2 Capacity Block in US East regions, falling to $37.76 in Europe (London) and Europe (Stockholm). On the ownership side, an NVIDIA HGX H100 8-GPU node carries a list price of $300,000 and draws 10.2 kW of rack power. This guide stacks both cost models with published rates, the discount levers that move them, and a checklist you can run against your own workload before signing a purchase order or a reservation.
The Cloud Side Numbers
Start from published rates, not from a vendor conversation, because list prices anchor every negotiation. The current Amazon EC2 Capacity Blocks for ML table prices a p5.48xlarge — 8 x H100 accelerators — at $41.528 USD per instance-hour in US East (N. Virginia), US East (Ohio) and US West (Oregon), while Europe (London) and Europe (Stockholm) come in at $37.76 USD, or $4.720 per accelerator-hour. Capacity Blocks charge an upfront reservation fee plus a per-second Linux OS fee of $0.000, which makes them a cleaner comparator against owned hardware than on-demand, where you also carry the risk of InsufficientCapacity errors during land-grab training windows.
Hyperscaler rates are the ceiling, not the market. As of the September 2026 review, tracked US on-demand H100 rates run from $2.99 per GPU-hour at Vultr to $10.98 on Google Cloud’s a3-highgpu-8g, with AWS normalized at $6.88 per GPU-hour in its cheapest US region and Azure at $6.98. Specialist providers such as Lambda ($3.29), Runpod ($3.49) and Nebius ($3.85) sit in the middle. That spread is the single largest controllable line in a cloud GPU budget, and it is why GPU total cost of ownership comparisons fail when they price only one vendor.
Discount levers compound on top of the headline rate. Google Cloud documents that Spot prices provide discounts of 60-91% off the corresponding on-demand price for most machine types and GPUs, and states that GPUs attached to standard VMs are eligible for resource-based committed use discounts when you create and attach a reservation at purchase time. Spot is a poor fit for multi-week training runs with checkpoint intolerance, but it is excellent for batch inference, hyperparameter sweeps and evaluation harnesses. If you are weighing the hyperscalers first, our GCP vs AWS pricing breakdown covers the non-GPU lines that quietly double a bill.
What On-Prem Actually Costs
Hardware acquisition is the visible part of the iceberg. Analyst tracking for July 2026 puts it plainly: a single NVIDIA H100 in the 80GB HBM3 Hopper configuration sells for approximately $25,000 to $40,000 depending on variant and reseller, with SXM5 modules at the top of that band and PCIe cards below. NVIDIA publishes no official price list, so treat any single quote as one data point in an opaque distribution shaped by volume, region and bundled services.
At node level, the reference 8U system that Dell, HPE, Supermicro and Lenovo all build around makes the math concrete: the HGX H100 8-GPU node carries a $300,000 list price, draws 10.2 kW per rack, and — amortized over 3 years with power priced at $0.05/kWh — works out to roughly $1,042 per GPU per month before a single engineer touches it. That $0.05/kWh assumption is optimistic for most of Europe, where commercial contracts often land higher; recompute it with your own tariff. Add a PUE estimate of 1.3 for an air-cooled hall, so every kilowatt-hour of compute becomes roughly 1.3 kWh of facility draw, then add the costs nobody lists on a datasheet: NVMe refresh, InfiniBand or 400 Gbps Ethernet switching, tenancy of an 8U/140 kg block, spares, and the platform team that patches drivers and Kubernetes GPU schedulers at 2 a.m.
Depreciation is the line most on-prem business cases understate. A 3-year amortization matches the support contract, but Hopper-class silicon is already being displaced by Blackwell systems whose full racks rent at materially higher rates because inference demand keeps repricing the installed base. Plan the exit value of your GPUs at close to zero, and the entrance price of the next generation at today’s, not yesterday’s, levels.
Utilization Decides the Winner
The break-even between renting and owning is almost entirely a utilization question, because cloud prices per hour while hardware prices per unit. A cluster that genuinely runs at 70%+ GPU utilization, week after week, spends most of its life in the region where owned hardware amortizes below rental rates. A cluster that bursts to 90% for two days and idles at 15% for the rest of the sprint is donating money to whoever owns the silicon.
The table below is an illustrative example — my own arithmetic with the assumptions stated, not a measured case study: one 8 x H100 workload, 70% sustained utilization, 18,396 GPU-busy hours over 3 years, compared at AWS’s Europe Capacity Block rate versus an owned HGX node kept powered around the clock.
| Cost line | Cloud Capacity Block (Europe) | Owned HGX H100 node |
|---|---|---|
| Compute | $37.76/hr × 18,396 h ≈ $694,600 | $300,000 hardware, 3-year write-off |
| Power | Included in the rate | 10.2 kW × 26,280 h ≈ 268,000 kWh at your tariff |
| Cooling and facility | Included | PUE estimate 1.3 on top of IT power |
| Operating system | $0 for Linux | Support contract, priced separately |
| People | Cloud platform overhead | Dedicated cluster operations team |
| Refresh | Switch SKU, no capital loss | New capital cycle at generation N+1 |
Read the two columns against each other, not in isolation. The cloud column scales to zero when the project ends; the owned column keeps drawing 10.2 kW whether the queue is full or empty. Once the workload does settle on one side, cost allocation tags that reconcile keep every GPU-hour attributable to the team that burned it — without attribution, utilization discipline decays within a quarter.
A Practical TCO Checklist
- Extract real GPU utilization from the last 90 days of billing and metrics data, not from roadmap hopes; annualize it.
- Price the owned side with your actual power contract, PUE, spares, networking and 3-year amortization — then add 20% for the lines you forgot.
- Compare against at least three cloud rates: a hyperscaler reservation, a Capacity Block and a specialist provider, per GPU-hour in one region set.
- Split the workload portfolio: Spot for interruptible jobs, committed rates or Capacity Blocks for long training, on-demand only for spillover.
- Model egress, storage and checkpoint I/O explicitly; they are the lines that surprise teams who compared only hourly compute.
- Revisit the break-even every quarter — provider rates move on supply and demand, and hardware resale values move with them.