Contents
Tags

An RTX 5090 launched at $1,999. In September 2026 the median US street price is around $4,699, and several retailers are past $5,000. That is not a scalping story that will pass in a fortnight — memory is now more than 80% of the bill of materials on a high-end card, GDDR7 contract prices have roughly tripled, and 2026 is the first year in about three decades in which Nvidia has not launched a new consumer architecture.
For local AI this changes the answer to the only question that matters. You do not buy a GPU for local inference — you buy memory, at some price per gigabyte, with a bandwidth figure attached. When the price per gigabyte moves by 2.4x, every buying guide written before this summer is wrong, including the parts of ours that quote list prices.
Runyard ranks hardware on dollars per usable gigabyte, because capacity is what decides whether a model loads at all. Usable memory is roughly 90% of nameplate on a discrete card — drivers, the desktop and fragmentation take the rest. Here is the 5090 at three prices:
Nothing about the card changed. It has the same 32 GB and the same 1,792 GB/s. Only the denominator moved, and it moved enough to take the 5090 from mid-table on value to the worst card in our catalogue on that measure.
At $4,699 the 5090 costs the same as a MacBook Pro M4 Max with 128 GB of unified memory. That is not a rhetorical flourish; it is the same number. What the two give you for it is very different:
Run our own arithmetic on what each holds at Q4_K_M with an 8K context. The 5090 tops out at a 32B model, needing 21.8 GB, and generates at roughly 80 tokens per second. The Mac reaches a 123B model, needing 75.6 GB, at roughly 6 tokens per second.
That is the entire trade, and it is a real one rather than a winner. Three times the capacity for a third of the speed. If your work is a 27B or 32B coding model and you want it to feel instant, the card is still right. If your work is a 70B or 120B model and you would rather wait than not run it at all, the Mac is now the cheaper way to get there — which was not true at list prices.
Apple's memory is not priced off the GDDR7 spot market, so unified-memory machines have held their pricing while discrete cards have not. That gap is the single biggest change in local AI hardware this year.
The cards nobody was excited about are the ones worth buying. Computed on our catalogue at current list prices:
The pattern is consistent: the further you get from the halo product, the less of the memory premium you are paying. A 16 GB card at $429 buys a gigabyte for less than a fifth of what the flagship charges.
A used RTX 3090 has 24 GB and 936 GB/s. At $900 that is $41.7 per gigabyte and about 49 tokens per second on a 27B model at Q4_K_M — within a few percent of what a 4090 does on the same model, for a fraction of the price. Even at $1,499 it matches the 5090's launch-price value while being available today.
Two things this arithmetic cannot check. Memory is what degrades under sustained load, and memory is exactly what you are buying it for, so run a memory stress test and a long generation session before paying. And a 3090 draws 350W while doing it.
Check any listing against every alternative that runs the same model, ranked on price per usable gigabyte.
Open the Used GPU Value Checker →Runyard's catalogue carries list prices, which means our price-per-gigabyte tool currently understates what a 5090 really costs you. We are treating that as a bug rather than a rounding error, and the calculators let you type the price you were actually quoted — which, in this market, is the only figure worth trusting.
Every card in the catalogue ranked by cost per usable gigabyte, with your own prices.
Open the price-per-GB tool →Tools
Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.
Newsletter