RRunyard Index
Runyard / GPU Index / Reference / Compute prices
Reference

What GPU compute costs

Every GPU in the index on one page — median, cheapest, spot, reserved, per month, memory per dollar — from 52 providers' own pricing as of 2026-09-18. Then the cheapest way to run each open model, and what changed lately. Bookmark it; it rewrites itself every six hours.

54 snapshots · 1085 configurations priced today

H100 SXM median
$3.50/hr
B200 median
$6.59/hr
Best memory per $
RTX 309096 GB per $/hr
Providers
52

Datacenter GPUs

GPUMemoryMedianCheapestatSpotReservedMonthGB / $·hrProviders
NVIDIANVIDIA H100 SXM80 GB$3.50$1.68Latitude.sh$2.04$2.65$2,51622.946
NVIDIANVIDIA A100 80GB80 GB$1.79$0.43Vast.ai$0.88$1.30$1,28944.737
NVIDIANVIDIA H200141 GB$4.45$2.00io.net$2.72$3.89$3,20431.735
NVIDIANVIDIA L40S48 GB$1.47$0.69Vast.ai$0.82$1.29$1,05832.730
NVIDIANVIDIA B200180 GB$6.59$3.75Packet.ai$4.02$6.55$4,74527.326
NVIDIANVIDIA L424 GB$0.80$0.32GPU.ai$5753019
NVIDIANVIDIA B300288 GB$7.85$4.50io.net$4.30$5,65236.717
NVIDIANVIDIA A100 40GB40 GB$1.48$0.79io.net$0.74$1,0662716
NVIDIANVIDIA L4048 GB$0.86$0.66io.net$0.80$0.70$61955.811
NVIDIANVIDIA V100 16GB16 GB$0.79$0.12Vast.ai$0.57$56920.311
NVIDIANVIDIA H100 PCIe80 GB$2.88$1.59io.net$2.00$2,07127.89
NVIDIANVIDIA T416 GB$0.59$0.15Akash$0.26$42527.19
AMDAMD MI300X192 GB$2.59$0.50RunPod Community Cloud$1.99$1,86574.17
NVIDIANVIDIA A1024 GB$1.29$1.00OVHcloud$0.60$92918.67
NVIDIANVIDIA A4048 GB$0.65$0.35RunPod Community Cloud$46873.87
NVIDIANVIDIA H100 NVL94 GB$3.03$2.59RunPod Community Cloud$1.29$2,183316
NVIDIANVIDIA H200 NVL141 GB$3.72$0.50RunPod Community Cloud$2,67837.95
NVIDIANVIDIA GH20096 GB$3.02$1.67Vultr$2.00$2,17431.85
NVIDIANVIDIA V100 32GB32 GB$2.66$1.57io.net$0.51$1,917124
NVIDIANVIDIA A10G24 GB$1.83$1.01Amazon Web Services$1,31813.13
AMDAMD MI325X256 GB$3.20$2.59Vultr$2.30$2,30080.12
AMDAMD MI355X288 GB$8.60$8.60Oracle Cloud$6,19233.51

Workstation GPUs

GPUMemoryMedianCheapestatSpotReservedMonthGB / $·hrProviders
NVIDIANVIDIA RTX PRO 6000 Blackwell96 GB$2.05$0.80HyperAI$1.22$1.30$1,47246.924
NVIDIANVIDIA RTX A600048 GB$0.56$0.33RunPod Community Cloud$0.40$0.35$40385.716
NVIDIANVIDIA RTX 6000 Ada48 GB$0.82$0.63Vast.ai$58758.910
NVIDIANVIDIA RTX A500024 GB$0.27$0.16RunPod Community Cloud$19488.98
NVIDIANVIDIA RTX A400016 GB$0.17$0.07Vast.ai$0.11$12294.17

Consumer GPUs

GPUMemoryMedianCheapestatSpotReservedMonthGB / $·hrProviders
NVIDIANVIDIA RTX 409024 GB$0.60$0.30io.net$0.16$4324013
NVIDIANVIDIA RTX 509032 GB$0.69$0.47GPU.ai$0.25$49746.49
NVIDIANVIDIA RTX 309024 GB$0.25$0.12Vast.ai$0.09$180967
NVIDIANVIDIA RTX 408016 GB$0.31$0.15Akash$0.11$22052.56
NVIDIANVIDIA RTX 508016 GB$0.53$0.26SaladCloud$0.15$38529.95

Median = median of each provider's own on-demand median. Cheapest = lowest single in-stock on-demand listing. Spot and reserved are medians where providers publish them. Month = median × 720 h.

Cheapest GPU for each model

For each reference model at 4-bit and 16-bit: the GPU whose median makes the cheapest set that holds the weights, KV cache at 8K context and overhead — one card where it fits, several where it does not. Priced at today's median and at the single cheapest listing.

ModelBitsCheapest set$/hr at medianAt cheapest listing~tok/s
Kimi K34-bit74× RTX 3090$18.50$8.61 Vast.ai
Kimi K316-bit260× RTX 3090$65.00$30.24 Vast.ai
DeepSeek V4.1-Flash4-bit22× RTX A4000$3.74$1.64 Vast.ai
DeepSeek V4.1-Flash16-bit52× RTX 3090$13.00$6.05 Vast.ai
Llama 4 Maverick4-bit16× RTX A4000$2.72$1.19 Vast.ai
Llama 4 Maverick16-bit38× RTX 3090$9.50$4.42 Vast.ai
gpt-oss-120b4-bit5× RTX A4000$0.85$0.37 Vast.ai
gpt-oss-120b16-bit11× RTX 3090$2.75$1.28 Vast.ai
Llama 3.1 70B4-bit4× RTX A4000$0.68$0.30 Vast.ai
Llama 3.1 70B16-bit7× RTX 3090$1.75$0.81 Vast.ai
Gemma 4 31B4-bit1× RTX 3090$0.25$0.12 Vast.ai43
Gemma 4 31B16-bit5× RTX A4000$0.85$0.37 Vast.ai
Qwen3 Coder 30B-A3B4-bit1× RTX 3090$0.25$0.12 Vast.ai403
Qwen3 Coder 30B-A3B16-bit3× RTX 3090$0.75$0.35 Vast.ai
Qwen3.8 27B4-bit1× RTX 3090$0.25$0.12 Vast.ai49
Qwen3.8 27B16-bit4× RTX A4000$0.68$0.30 Vast.ai
gpt-oss-20b4-bit1× RTX A4000$0.17$0.07 Vast.ai177
gpt-oss-20b16-bit4× RTX A4000$0.68$0.30 Vast.ai
Llama 3.1 8B4-bit1× RTX A4000$0.17$0.07 Vast.ai80
Llama 3.1 8B16-bit1× RTX 3090$0.25$0.12 Vast.ai47

tok/s is single-stream decode from memory bandwidth, derated 20% — an estimate for ranking, not a benchmark. Multi-GPU sets have no tok/s figure.

Recent price changes

WhenProviderGPUFromToChange
2026-09-18 22:13Vast.aiRTX 3090$0.21$0.27+29.3%
2026-09-18 22:13Vast.aiRTX 5080$0.74$0.53-27.6%
2026-09-18 22:13SpheronRTX 4090$0.58$0.72+24.1%
2026-09-18 22:13Vast.aiRTX 4090$0.60$0.72+19.1%
2026-09-18 22:13Vast.aiB200$7.38$8.75+18.7%
2026-09-18 22:13Vast.aiRTX 5090$0.60$0.68+12.2%
2026-09-18 22:13GPU.aiRTX 4090$0.54$0.60+11.1%
2026-09-18 22:13GPU.aiA100 80GB$1.10$1.21+10.5%
2026-09-18 22:13GPU.aiB200$7.09$7.78+9.7%
2026-09-18 22:13Vast.aiH100 SXM$3.79$3.45-9.1%
2026-09-18 22:13Vast.aiRTX A5000$0.23$0.22-6.6%
2026-09-18 22:13SpheronB300$9.62$10.21+6.1%
2026-09-18 22:13GPU.aiRTX A5000$0.26$0.27+5.9%
2026-09-18 22:13GPU.aiRTX 3090$0.32$0.34+4.7%
2026-09-18 22:13GPU.aiH200 NVL$3.62$3.46-4.4%

Questions

How much does GPU compute cost in 2026?

As of 2026-09-18, across 52 providers: an H100 SXM rents for a median of $3.50 per GPU-hour (cheapest $1.68), a B200 for $6.59, an A100 80GB for $1.79, and a consumer RTX 4090 for $0.60. Multiply by 720 for a month. The table above has every GPU we track.

What is the cheapest GPU per gigabyte of memory?

The NVIDIA RTX 3090: 96 GB of memory per dollar-hour at its median of $0.25. Memory per dollar is the right lens when the job is holding a model, not training one — most inference is memory-bound.

How much does it cost to generate a million tokens?

It depends on the model and the card. Llama 3.1 70B at 4-bit on one H100 SXM: about $14.28 per million generated tokens at today's median, single stream, from memory bandwidth. The cheapest set that holds it is 4× RTX A4000 at $0.68/hr. Batching many requests lowers the per-token cost by an order of magnitude; these figures are the one-user floor, useful for ranking cards against each other.

Why do the same GPU's prices differ 5× between providers?

Three markets share one name. Hyperscalers price the instance — CPU, RAM, network, ecosystem — and the GPU is a fraction of it. Neoclouds price the GPU. Marketplaces are independent hosts competing on price with consumer-grade guarantees. The median sits with the neoclouds; the cheapest listing is nearly always a marketplace host; the most expensive is a hyperscaler. Each is the right answer for a different buyer.

How is this table produced?

Every provider's public pricing is read every six hours. Each priced unit becomes a row — instance, SKU, offer — divided by its GPU count. Per GPU, each provider is reduced to its own median, and the headline is the median of those. Spot and reserved are tracked the same way but never enter the headline. There are 54 snapshots so far; every one is kept in a public repository. The full method is on the methodology page.

How every number on this page is produced, in full.
Methodology →