Two GPUs, or two providers, side by side — from the same index, priced 2026-09-18. GPU pages compare rental price, memory, bandwidth, which models fit on one card and what a million generated tokens costs on each. Provider pages compare the two on every GPU both sell.
Every pair among the widest-coverage providers that share at least three GPUs.
On 2026-09-18 the H100 SXM median is $3.50 per GPU-hour and the A100 80GB $1.79. Both hold 80 GB, so the same models fit on one card; the H100's memory bandwidth is 3,350 GB/s against 2,039, so it generates tokens roughly 1.6× faster per card. If the H100 costs less than 1.6× the A100 where you rent, it is the better per-token deal; the comparison page prices this per model.
H200 median $4.45, B200 $6.59. The B200 has 180 GB against 141 and 8,000 GB/s against 4,800, so it holds larger models on one card and streams them faster. For a model that already fits on an H200, the question is whether the B200's bandwidth premium is smaller than its price premium — the page computes cost per million tokens on each.
RTX 4090 median $0.60, RTX 5090 $0.69. The 5090's 32 GB against 24 lets a 4-bit 30B model with room for context fit where the 4090 is tight, and its 1,792 GB/s against 1,008 is a large speed step. Both are consumer cards mostly on marketplaces, so availability and host quality vary.
Every pair among the widest-coverage providers that share at least three GPUs gets a page: each GPU both sell, with both prices, who is cheaper on each, and the market median beside them. Only GPUs both actually price appear, so it is a like-for-like table, not a marketing sheet. Pairs are rebuilt from each snapshot.
The GPU's hourly median divided by an estimate of how many tokens it generates in an hour for that model, single stream, from memory bandwidth — every token streams the active weights once. It is a floor, not a benchmark, and batching lowers it by an order of magnitude; it exists to rank cards against each other for the same model.