โ† All articles

GPU cloud pricing compared: what an H100 really costs per hour

By LW Forge โ€” maintainer of LLM Scout ยท Updated August 24, 2026

The same NVIDIA H100 costs $1.65 an hour on a marketplace and $11.06 an hour on Google Cloud. Identical silicon, a 6.7ร— spread, and no amount of benchmarking will close it โ€” you're paying for contract structure, support, network fabric and who eats the risk when a node dies. Before you decide whether to self-host inference, you need to know which of those you're actually buying.

What an H100 costs, by provider

These are on-demand, per-GPU-hour rates checked on 2026-07-06. Unlike our LLM API prices, GPU figures on this site are a hand-researched snapshot, not a live feed โ€” there is no free public API that aggregates GPU cloud pricing. Always open the provider's page before committing budget; our sources page lists them.

ProviderType$/hr~$/month (730h)
Vast.aiMarketplace$1.65*$1,205
RunPodGPU cloud$2.89$2,110
Lambda LabsGPU cloud$3.99$2,913
CoreWeaveGPU cloud$4.25*$3,103
Paperspace / DigitalOceanGPU cloud$5.95$4,344
AWS (P5)Hyperscaler$6.88$5,022
Azure (NC H100 v5)Hyperscaler$6.98$5,095
Google Cloud (A3)Hyperscaler$11.06$8,074

* Estimates. Vast.ai is a marketplace where verified hosts ranged roughly $1.50โ€“$1.87/hr; CoreWeave's per-GPU Classic rate is derived rather than published as a single line item. Treat both as a range, not a quote.

The A100 80GB tells the same story more cheaply: Vast $0.90, RunPod $1.39, CoreWeave $2.21, Lambda $2.79, Paperspace $3.18, Azure $3.67, AWS $4.10. And if your model fits in less VRAM, the floor drops fast โ€” an L40S runs $0.99/hr on RunPod, an RTX 4090 $0.69 there or about $0.35 on Vast. Most inference workloads people describe as "needing an H100" fit comfortably on a 48GB card.

Three provider categories explain the spread. Hyperscalers (AWS, GCP, Azure) charge 2โ€“4ร— the neocloud rate and you get enterprise contracts, compliance paperwork, private networking and integration with the rest of your stack. Specialized GPU clouds (RunPod, Lambda, CoreWeave, Paperspace) strip that down to a container and a public IP. Marketplaces (Vast.ai) resell other people's idle hardware โ€” cheapest per hour, most variable in reliability, disk speed and uptime. For batch jobs that can be restarted, that variance is nearly free. For a user-facing endpoint, it isn't.

A spot check on 2026-08-04 suggests the market has drifted down since our snapshot: neoclouds now advertise H100 from roughly $2.00โ€“$2.50/hr, and B200 rates run about $3.70โ€“$7.00 per GPU-hour outside the hyperscalers. Direction of travel: cheaper.

Spot, on-demand, reserved

On-demand is the number above: start it, stop it, pay by the second or minute. Spot / interruptible / community tiers typically run 40โ€“60% below on-demand and can be reclaimed with little warning โ€” RunPod's Community Cloud 4090 at about $0.34 versus $0.69 on Secure Cloud is a typical gap. Reserved commitments of one to three years usually land 30โ€“40% under on-demand, which only pays off if your utilization is genuinely continuous. The trap is committing to a year of capacity for a workload whose traffic hasn't stabilized yet.

The only calculation that matters

Hourly price is the wrong unit. What you actually want is cost per million tokens served, compared against just paying an API. Run one H100 on RunPod at $2.89/hr, 24/7: $2,110 a month. Now spend that same money on API calls:

ModelOutput $/1MOutput tokens $2,110 buys
Llama 4 Maverick$0.60~3.5 billion
Nemotron 3 Ultra$2.20~960 million
Gemini 2.5 Flash$2.50~844 million
Claude Sonnet 5$15.00~141 million
GPT-5.6 Sol$30.00~70 million

To beat Maverick's rate, your H100 has to sustain about 1,340 output tokens per second, every second of the month. That's achievable for a small model under vLLM with heavy batching โ€” but only at near-full utilization. Drop to a realistic 25% (a product with a daily peak and a quiet night) and your effective cost is roughly $2.41 per million output tokens: four times Maverick's rate, and worse than Nemotron 3 Ultra's hosted price.

That's the whole decision, and it isn't about hardware. Self-hosting competes with the cheap open-weights tier, essentially never with a frontier model you couldn't host anyway. If you want to see how that plays out from the model side, Nemotron 3 and the open-weights math is the other half of this question. Price your own configuration in the GPU cloud cost calculator, and check your real token volumes in the token calculator before trusting a word-count estimate.

The costs that don't show up in the hourly rate

Idle time. You rent hours, not tokens. A pod at 20% utilization costs 5ร— per token what the spec sheet implies. This is the single largest hidden cost and nobody models it up front.

Storage and cold starts. Weights live on a persistent volume at roughly $0.05โ€“$0.10/GB/month. A 550B checkpoint in FP8 is a few hundred gigabytes โ€” tens of dollars a month, plus minutes of pull time on every cold start, which quietly kills any plan to scale to zero.

Egress. Text tokens are negligible: a billion output tokens is about 4GB, under a dollar at hyperscaler rates. The bill arrives when you move checkpoints โ€” 500GB out of AWS at $0.09/GB is $45 per copy, every time you switch providers.

Ops. Someone runs the server, tunes batching, debugs OOMs and handles dead nodes. Four hours a week of an engineer at a $80/hr loaded cost is about $1,386 a month โ€” more than the H100 itself at RunPod prices.

Add those up and the honest rule is this: self-host when you have continuous high volume, a hard data-residency requirement, or a fine-tuned model with no hosted equivalent. Otherwise the API wins on total cost, and the argument was never really about the GPU. If your motivation is purely price, start with the cheapest LLM APIs of 2026 or the DeepSeek cost comparison โ€” the ceiling on savings there is usually higher than anything a rented H100 gets you.