RRunyard Index
Runyard / GPU Index / NVIDIA H100 SXM
NVIDIA
Rental price index

NVIDIA H100 SXM

The reference datacenter GPU. Underlies the CME H100 rental index.

Prices collected 2026-09-12 · 23 providers, 62 configurations

Compare vs other GPUs →
On-demandOn-demand (archive, Lambda only)Spot
$0.00$1.25$2.50$3.75$5.00Sep 11 08:15Sep 11 10:33Sep 12 04:39Sep 12 20:35

Index median price per GPU per hour, by billing type. Nothing is interpolated.

At a glance

Median on-demand
$3.49/GPU/hr
Cheapest on-demand
$1.68Latitude.sh
Floor, any billing
$1.68incl. spot
Per month at 720h
$2,513

NVIDIA H100 SXM pricing by provider

One card per provider, cheapest first. A configuration is one priced unit — an instance type, a SKU, a marketplace offer — and the per-GPU price is that unit's price divided by its GPU count. Open a card to see every configuration with its vCPU, RAM and region where the provider publishes them.

23 providers · 57 configurations
HyperAIIn stock
China1 config1xcollected 2026-09-12
On-Demand from $1.80
From
$1.80 / GPU / hr
On-DemandVisit website →
SpheronIn stock
2 configs1xcollected 2026-09-12
On-Demand from $2.01
From
$2.01 / GPU / hr
On-DemandVisit website →
Vast.aiIn stock
USA8 configs1x–2xcollected 2026-09-12
On-Demand from $2.27
From
$2.27 / GPU / hr
On-Demand · The Netherlands, NLVisit website →
OblivusIn stock
UK6 configs1x–8xcollected 2026-09-12
On-Demand from $2.49
From
$2.49 / GPU / hr
On-DemandVisit website →
LyceumIn stock
Germany3 configs1xcollected 2026-09-12
On-Demand from $2.49
From
$2.49 / GPU / hr
On-DemandVisit website →
KoyebIn stock
France4 configs1x–8xcollected 2026-09-12
On-Demand from $2.50
From
$2.50 / GPU / hr
On-DemandVisit website →
RuncrateIn stock
1 config1xcollected 2026-09-12
On-Demand from $2.75
From
$2.75 / GPU / hr
On-DemandVisit website →
CivoIn stock
UK2 configs1x–8xcollected 2026-09-12
On-Demand from $2.99
From
$2.99 / GPU / hr
On-DemandVisit website →
GPU.aiIn stock
UAE2 configs1xcollected 2026-09-12
On-Demand from $3.06
From
$3.06 / GPU / hr
On-DemandVisit website →
BeamIn stock
USA1 config1xcollected 2026-09-12
On-Demand from $3.50
From
$3.50 / GPU / hr
On-DemandVisit website →
LambdaIn stock
USA4 configs1x–8xcollected 2026-09-12
On-Demand from $3.99
From
$3.99 / GPU / hr
On-DemandVisit website →
fal.aiIn stock
USA1 config1xcollected 2026-09-12
On-Demand from $4.50
From
$4.50 / GPU / hr
On-DemandVisit website →
ReplicateIn stock
USA4 configs1x–8xcollected 2026-09-12
On-Demand from $5.49
From
$5.49 / GPU / hr
On-DemandVisit website →
CoreWeaveIn stock
USA1 config8xcollected 2026-09-12
On-Demand from $6.16
From
$6.16 / GPU / hr
On-DemandVisit website →
USA5 configs8xcollected 2026-09-12
On-Demand from $11.06Spot from $2.04
From
$11.06 / GPU / hr
On-Demand · eastusVisit website →

What fits on a NVIDIA H100 SXM

Weights + KV cache at 8K context + runtime overhead, against 72 GB usable of the 80 GB on the card. Capacity follows total parameters, even for mixture-of-experts models — any expert may be needed next. Where one card is not enough, the table says how many are, and prices the set at today's median.

ModelParams4-bit8-bit16-bit
Kimi K32.8T104B active22× GPU$76.78/hr42× GPU$146.58/hr78× GPU$272.22/hr
DeepSeek V4.1-Flash552B16B active5× GPU$17.45/hr9× GPU$31.41/hr16× GPU$55.84/hr
Llama 4 Maverick400B17B active4× GPU$13.96/hr6× GPU$20.94/hr12× GPU$41.88/hr
gpt-oss-120b117B5.1B active✓ 67.9 GB~934 tok/s2× GPU$6.98/hr4× GPU$13.96/hr
Llama 3.1 70B70B✓ 44.5 GB~68 tok/s2× GPU$6.98/hr3× GPU$10.47/hr
Gemma 4 31B31B✓ 21.2 GB~154 tok/s✓ 36.7 GB~81 tok/s✓ 65.7 GB~43 tok/s
Qwen3 Coder 30B-A3B30.5B3.3B active✓ 19 GB~1444 tok/s✓ 34.3 GB~764 tok/s✓ 62.9 GB~406 tok/s
Qwen3.8 27B27B✓ 18.7 GB~176 tok/s✓ 32.2 GB~93 tok/s✓ 57.6 GB~50 tok/s
gpt-oss-20b21B3.6B active✓ 13.7 GB~1323 tok/s✓ 24.2 GB~701 tok/s✓ 43.9 GB~372 tok/s
Llama 3.1 8B8B✓ 6.9 GB~596 tok/s✓ 10.9 GB~315 tok/s✓ 18.4 GB~168 tok/s

tok/s is decode throughput from memory bandwidth, derated 20%, single stream — an estimate, not a benchmark. Set cost uses today's median × GPUs required.

Specifications

Memory80 GB HBM3
Memory bandwidth3,350 GB/s
ArchitectureHopper
VendorNVIDIA
ClassDatacenter
Released2022-09

Bandwidth is listed because it governs generation speed: each token streams the active weights once, so tokens per second is bandwidth divided by bytes read. TFLOPs decide prompt processing and training, not how fast text appears.

Alternatives

Questions

How much does it cost to rent a NVIDIA H100 SXM?

As of 2026-09-12, the median on-demand price across 23 providers is $3.49 per GPU-hour, about $2,513 a month at 720 hours. The cheapest we saw was $1.68 at Latitude.sh.

Why do prices for the same GPU vary so much?

Hyperscalers bundle CPU, RAM and network into an instance price and charge for the ecosystem around it. Neoclouds sell the GPU more directly. Marketplaces are independent hosts competing on price, with the trade-offs of shared, third-party hardware. The per-GPU figure is the instance price divided by GPU count, which is the only way to compare an 8-GPU node with a single-card listing.

Is the median the price I will pay?

No — it is the middle of the market. Each provider's own median is taken first so a provider with many instance sizes counts once, then the median across providers. Spot and reserved rates are excluded from it; they set the floor shown separately. Use it to judge whether a quote is high or low, then verify with the provider.

Which models can a NVIDIA H100 SXM run?

With 80 GB per GPU, the table above shows which reference models fit on one card at 4, 8 and 16-bit, using the same memory arithmetic as the rest of Runyard: weights plus KV cache plus runtime overhead, with capacity following total parameters even for mixture-of-experts models. Where a model needs more than one GPU it says how many and what that set costs at today's median.

Own the hardware instead? Work out what your card holds with the same arithmetic.
Check what fits →