Every GPU in the index on one page — median, cheapest, spot, reserved, per month, memory per dollar — from 52 providers' own pricing as of 2026-09-18. Then the cheapest way to run each open model, and what changed lately. Bookmark it; it rewrites itself every six hours.
54 snapshots · 1085 configurations priced today
| GPU | Memory | Median | Cheapest | at | Spot | Reserved | Month | GB / $·hr | Providers |
|---|---|---|---|---|---|---|---|---|---|
| 80 GB | $3.50 | $1.68 | $2.04 | $2.65 | $2,516 | 22.9 | 46 | ||
| 80 GB | $1.79 | $0.43 | $0.88 | $1.30 | $1,289 | 44.7 | 37 | ||
| 141 GB | $4.45 | $2.00 | $2.72 | $3.89 | $3,204 | 31.7 | 35 | ||
| 48 GB | $1.47 | $0.69 | $0.82 | $1.29 | $1,058 | 32.7 | 30 | ||
| 180 GB | $6.59 | $3.75 | $4.02 | $6.55 | $4,745 | 27.3 | 26 | ||
| 24 GB | $0.80 | $0.32 | — | — | $575 | 30 | 19 | ||
| 288 GB | $7.85 | $4.50 | $4.30 | — | $5,652 | 36.7 | 17 | ||
| 40 GB | $1.48 | $0.79 | $0.74 | — | $1,066 | 27 | 16 | ||
| 48 GB | $0.86 | $0.66 | $0.80 | $0.70 | $619 | 55.8 | 11 | ||
| 16 GB | $0.79 | $0.12 | $0.57 | — | $569 | 20.3 | 11 | ||
| 80 GB | $2.88 | $1.59 | $2.00 | — | $2,071 | 27.8 | 9 | ||
| 16 GB | $0.59 | $0.15 | $0.26 | — | $425 | 27.1 | 9 | ||
| 192 GB | $2.59 | $0.50 | — | $1.99 | $1,865 | 74.1 | 7 | ||
| 24 GB | $1.29 | $1.00 | $0.60 | — | $929 | 18.6 | 7 | ||
| 48 GB | $0.65 | $0.35 | — | — | $468 | 73.8 | 7 | ||
| 94 GB | $3.03 | $2.59 | $1.29 | — | $2,183 | 31 | 6 | ||
| 141 GB | $3.72 | $0.50 | — | — | $2,678 | 37.9 | 5 | ||
| 96 GB | $3.02 | $1.67 | — | $2.00 | $2,174 | 31.8 | 5 | ||
| 32 GB | $2.66 | $1.57 | $0.51 | — | $1,917 | 12 | 4 | ||
| 24 GB | $1.83 | $1.01 | — | — | $1,318 | 13.1 | 3 | ||
| 256 GB | $3.20 | $2.59 | — | $2.30 | $2,300 | 80.1 | 2 | ||
| 288 GB | $8.60 | $8.60 | — | — | $6,192 | 33.5 | 1 |
| GPU | Memory | Median | Cheapest | at | Spot | Reserved | Month | GB / $·hr | Providers |
|---|---|---|---|---|---|---|---|---|---|
| 96 GB | $2.05 | $0.80 | $1.22 | $1.30 | $1,472 | 46.9 | 24 | ||
| 48 GB | $0.56 | $0.33 | $0.40 | $0.35 | $403 | 85.7 | 16 | ||
| 48 GB | $0.82 | $0.63 | — | — | $587 | 58.9 | 10 | ||
| 24 GB | $0.27 | $0.16 | — | — | $194 | 88.9 | 8 | ||
| 16 GB | $0.17 | $0.07 | — | $0.11 | $122 | 94.1 | 7 |
| GPU | Memory | Median | Cheapest | at | Spot | Reserved | Month | GB / $·hr | Providers |
|---|---|---|---|---|---|---|---|---|---|
| 24 GB | $0.60 | $0.30 | $0.16 | — | $432 | 40 | 13 | ||
| 32 GB | $0.69 | $0.47 | $0.25 | — | $497 | 46.4 | 9 | ||
| 24 GB | $0.25 | $0.12 | $0.09 | — | $180 | 96 | 7 | ||
| 16 GB | $0.31 | $0.15 | $0.11 | — | $220 | 52.5 | 6 | ||
| 16 GB | $0.53 | $0.26 | $0.15 | — | $385 | 29.9 | 5 |
Median = median of each provider's own on-demand median. Cheapest = lowest single in-stock on-demand listing. Spot and reserved are medians where providers publish them. Month = median × 720 h.
For each reference model at 4-bit and 16-bit: the GPU whose median makes the cheapest set that holds the weights, KV cache at 8K context and overhead — one card where it fits, several where it does not. Priced at today's median and at the single cheapest listing.
| Model | Bits | Cheapest set | $/hr at median | At cheapest listing | ~tok/s |
|---|---|---|---|---|---|
| Kimi K3 | 4-bit | 74× RTX 3090 | $18.50 | $8.61 Vast.ai | — |
| Kimi K3 | 16-bit | 260× RTX 3090 | $65.00 | $30.24 Vast.ai | — |
| DeepSeek V4.1-Flash | 4-bit | 22× RTX A4000 | $3.74 | $1.64 Vast.ai | — |
| DeepSeek V4.1-Flash | 16-bit | 52× RTX 3090 | $13.00 | $6.05 Vast.ai | — |
| Llama 4 Maverick | 4-bit | 16× RTX A4000 | $2.72 | $1.19 Vast.ai | — |
| Llama 4 Maverick | 16-bit | 38× RTX 3090 | $9.50 | $4.42 Vast.ai | — |
| gpt-oss-120b | 4-bit | 5× RTX A4000 | $0.85 | $0.37 Vast.ai | — |
| gpt-oss-120b | 16-bit | 11× RTX 3090 | $2.75 | $1.28 Vast.ai | — |
| Llama 3.1 70B | 4-bit | 4× RTX A4000 | $0.68 | $0.30 Vast.ai | — |
| Llama 3.1 70B | 16-bit | 7× RTX 3090 | $1.75 | $0.81 Vast.ai | — |
| Gemma 4 31B | 4-bit | 1× RTX 3090 | $0.25 | $0.12 Vast.ai | 43 |
| Gemma 4 31B | 16-bit | 5× RTX A4000 | $0.85 | $0.37 Vast.ai | — |
| Qwen3 Coder 30B-A3B | 4-bit | 1× RTX 3090 | $0.25 | $0.12 Vast.ai | 403 |
| Qwen3 Coder 30B-A3B | 16-bit | 3× RTX 3090 | $0.75 | $0.35 Vast.ai | — |
| Qwen3.8 27B | 4-bit | 1× RTX 3090 | $0.25 | $0.12 Vast.ai | 49 |
| Qwen3.8 27B | 16-bit | 4× RTX A4000 | $0.68 | $0.30 Vast.ai | — |
| gpt-oss-20b | 4-bit | 1× RTX A4000 | $0.17 | $0.07 Vast.ai | 177 |
| gpt-oss-20b | 16-bit | 4× RTX A4000 | $0.68 | $0.30 Vast.ai | — |
| Llama 3.1 8B | 4-bit | 1× RTX A4000 | $0.17 | $0.07 Vast.ai | 80 |
| Llama 3.1 8B | 16-bit | 1× RTX 3090 | $0.25 | $0.12 Vast.ai | 47 |
tok/s is single-stream decode from memory bandwidth, derated 20% — an estimate for ranking, not a benchmark. Multi-GPU sets have no tok/s figure.
| When | Provider | GPU | From | To | Change |
|---|---|---|---|---|---|
| 2026-09-18 22:13 | RTX 3090 | $0.21 | $0.27 | +29.3% | |
| 2026-09-18 22:13 | RTX 5080 | $0.74 | $0.53 | -27.6% | |
| 2026-09-18 22:13 | RTX 4090 | $0.58 | $0.72 | +24.1% | |
| 2026-09-18 22:13 | RTX 4090 | $0.60 | $0.72 | +19.1% | |
| 2026-09-18 22:13 | B200 | $7.38 | $8.75 | +18.7% | |
| 2026-09-18 22:13 | RTX 5090 | $0.60 | $0.68 | +12.2% | |
| 2026-09-18 22:13 | RTX 4090 | $0.54 | $0.60 | +11.1% | |
| 2026-09-18 22:13 | A100 80GB | $1.10 | $1.21 | +10.5% | |
| 2026-09-18 22:13 | B200 | $7.09 | $7.78 | +9.7% | |
| 2026-09-18 22:13 | H100 SXM | $3.79 | $3.45 | -9.1% | |
| 2026-09-18 22:13 | RTX A5000 | $0.23 | $0.22 | -6.6% | |
| 2026-09-18 22:13 | B300 | $9.62 | $10.21 | +6.1% | |
| 2026-09-18 22:13 | RTX A5000 | $0.26 | $0.27 | +5.9% | |
| 2026-09-18 22:13 | RTX 3090 | $0.32 | $0.34 | +4.7% | |
| 2026-09-18 22:13 | H200 NVL | $3.62 | $3.46 | -4.4% |
As of 2026-09-18, across 52 providers: an H100 SXM rents for a median of $3.50 per GPU-hour (cheapest $1.68), a B200 for $6.59, an A100 80GB for $1.79, and a consumer RTX 4090 for $0.60. Multiply by 720 for a month. The table above has every GPU we track.
The NVIDIA RTX 3090: 96 GB of memory per dollar-hour at its median of $0.25. Memory per dollar is the right lens when the job is holding a model, not training one — most inference is memory-bound.
It depends on the model and the card. Llama 3.1 70B at 4-bit on one H100 SXM: about $14.28 per million generated tokens at today's median, single stream, from memory bandwidth. The cheapest set that holds it is 4× RTX A4000 at $0.68/hr. Batching many requests lowers the per-token cost by an order of magnitude; these figures are the one-user floor, useful for ranking cards against each other.
Three markets share one name. Hyperscalers price the instance — CPU, RAM, network, ecosystem — and the GPU is a fraction of it. Neoclouds price the GPU. Marketplaces are independent hosts competing on price with consumer-grade guarantees. The median sits with the neoclouds; the cheapest listing is nearly always a marketplace host; the most expensive is a hyperscaler. Each is the right answer for a different buyer.
Every provider's public pricing is read every six hours. Each priced unit becomes a row — instance, SKU, offer — divided by its GPU count. Per GPU, each provider is reduced to its own median, and the headline is the median of those. Spot and reserved are tracked the same way but never enter the headline. There are 54 snapshots so far; every one is kept in a public repository. The full method is on the methodology page.