Runyard / GPUs

GPUs for local AI, compared

Two numbers decide a GPU’s worth for local inference: how much it holds, and how fast it reads. Neither appears in a marketing headline.

VRAM capacity determines whether a model runs at all. Memory bandwidth determines how quickly it generates once it does, because decoding a token means streaming the active weights out of memory. Cards with identical capacity routinely differ threefold in bandwidth, which is why two GPUs that look equivalent on a spec sheet can feel nothing alike in use.

GPUVRAMBandwidthPowerMSRPModels it runs
RTX 509032 GB1,792 GB/s575W$1,99953
RTX 409024 GB1,008 GB/s450W$1,59953
RX 7900 XTX24 GB960 GB/s355W$99953
RTX 309024 GB936 GB/s350W$1,49953
RTX 508016 GB960 GB/s360W$99938
RTX 5070 Ti16 GB896 GB/s300W$74938
RTX 4080 SUPER16 GB736 GB/s320W$99938
RTX 4070 Ti SUPER16 GB672 GB/s285W$79938
RX 9070 XT16 GB645 GB/s304W$59938
RTX 5060 Ti 16GB16 GB448 GB/s180W$42938
RTX 4060 Ti 16GB16 GB288 GB/s165W$44938
RTX 507012 GB672 GB/s250W$54926
Arc B58012 GB456 GB/s190W$24926

The final column counts how many of the models we track fit entirely in each card’s memory at some usable quantisation. It is the most honest single summary of a GPU’s usefulness for local AI, and it tracks price poorly — which is rather the point of publishing it.

Tools and other sections