🏢PCIe H200 for air-cooled servers.
Prices collected 2026-09-12 · 5 providers, 23 configurations
Compare vs other GPUs →Index median price per GPU per hour, by billing type. Nothing is interpolated.
One card per provider, cheapest first. A configuration is one priced unit — an instance type, a SKU, a marketplace offer — and the per-GPU price is that unit's price divided by its GPU count. Open a card to see every configuration with its vCPU, RAM and region where the provider publishes them.
Weights + KV cache at 8K context + runtime overhead, against 127 GB usable of the 141 GB on the card. Capacity follows total parameters, even for mixture-of-experts models — any expert may be needed next. Where one card is not enough, the table says how many are, and prices the set at today's median.
| Model | Params | 4-bit | 8-bit | 16-bit |
|---|---|---|---|---|
| 2.8T104B active | 13× GPU$48.36/hr | 24× GPU$89.28/hr | 45× GPU$167.40/hr | |
| 552B16B active | 3× GPU$11.16/hr | 5× GPU$18.60/hr | 9× GPU$33.48/hr | |
| 400B17B active | 2× GPU$7.44/hr | 4× GPU$14.88/hr | 7× GPU$26.04/hr | |
| 117B5.1B active | ✓ 67.9 GB~1339 tok/s | ✓ 126.4 GB~709 tok/s | 2× GPU$7.44/hr | |
| 70B | ✓ 44.5 GB~98 tok/s | ✓ 79.5 GB~52 tok/s | 2× GPU$7.44/hr | |
| 31B | ✓ 21.2 GB~220 tok/s | ✓ 36.7 GB~117 tok/s | ✓ 65.7 GB~62 tok/s | |
| 30.5B3.3B active | ✓ 19 GB~2069 tok/s | ✓ 34.3 GB~1095 tok/s | ✓ 62.9 GB~582 tok/s | |
| 27B | ✓ 18.7 GB~253 tok/s | ✓ 32.2 GB~134 tok/s | ✓ 57.6 GB~71 tok/s | |
| 21B3.6B active | ✓ 13.7 GB~1896 tok/s | ✓ 24.2 GB~1004 tok/s | ✓ 43.9 GB~533 tok/s | |
| 8B | ✓ 6.9 GB~853 tok/s | ✓ 10.9 GB~452 tok/s | ✓ 18.4 GB~240 tok/s |
tok/s is decode throughput from memory bandwidth, derated 20%, single stream — an estimate, not a benchmark. Set cost uses today's median × GPUs required.
Bandwidth is listed because it governs generation speed: each token streams the active weights once, so tokens per second is bandwidth divided by bytes read. TFLOPs decide prompt processing and training, not how fast text appears.
As of 2026-09-12, the median on-demand price across 5 providers is $3.72 per GPU-hour, about $2,678 a month at 720 hours. The cheapest we saw was $0.50 at RunPod Community Cloud.
Hyperscalers bundle CPU, RAM and network into an instance price and charge for the ecosystem around it. Neoclouds sell the GPU more directly. Marketplaces are independent hosts competing on price, with the trade-offs of shared, third-party hardware. The per-GPU figure is the instance price divided by GPU count, which is the only way to compare an 8-GPU node with a single-card listing.
No — it is the middle of the market. Each provider's own median is taken first so a provider with many instance sizes counts once, then the median across providers. Spot and reserved rates are excluded from it; they set the floor shown separately. Use it to judge whether a quote is high or low, then verify with the provider.
With 141 GB per GPU, the table above shows which reference models fit on one card at 4, 8 and 16-bit, using the same memory arithmetic as the rest of Runyard: weights plus KV cache plus runtime overhead, with capacity following total parameters even for mixture-of-experts models. Where a model needs more than one GPU it says how many and what that set costs at today's median.