Two numbers decide a GPU’s worth for local inference: how much it holds, and how fast it reads. Neither appears in a marketing headline.
VRAM capacity determines whether a model runs at all. Memory bandwidth determines how quickly it generates once it does, because decoding a token means streaming the active weights out of memory. Cards with identical capacity routinely differ threefold in bandwidth, which is why two GPUs that look equivalent on a spec sheet can feel nothing alike in use.
| GPU | VRAM | Bandwidth | Power | MSRP | Models it runs |
|---|---|---|---|---|---|
| RTX 5090 | 32 GB | 1,792 GB/s | 575W | $1,999 | 53 |
| RTX 4090 | 24 GB | 1,008 GB/s | 450W | $1,599 | 53 |
| RX 7900 XTX | 24 GB | 960 GB/s | 355W | $999 | 53 |
| RTX 3090 | 24 GB | 936 GB/s | 350W | $1,499 | 53 |
| RTX 5080 | 16 GB | 960 GB/s | 360W | $999 | 38 |
| RTX 5070 Ti | 16 GB | 896 GB/s | 300W | $749 | 38 |
| RTX 4080 SUPER | 16 GB | 736 GB/s | 320W | $999 | 38 |
| RTX 4070 Ti SUPER | 16 GB | 672 GB/s | 285W | $799 | 38 |
| RX 9070 XT | 16 GB | 645 GB/s | 304W | $599 | 38 |
| RTX 5060 Ti 16GB | 16 GB | 448 GB/s | 180W | $429 | 38 |
| RTX 4060 Ti 16GB | 16 GB | 288 GB/s | 165W | $449 | 38 |
| RTX 5070 | 12 GB | 672 GB/s | 250W | $549 | 26 |
| Arc B580 | 12 GB | 456 GB/s | 190W | $249 | 26 |
The final column counts how many of the models we track fit entirely in each card’s memory at some usable quantisation. It is the most honest single summary of a GPU’s usefulness for local AI, and it tracks price poorly — which is rather the point of publishing it.