For local AI the honest first filter is dollars per gigabyte, because memory decides what runs at all. Gaming value rankings do not transfer.

Computed from this tool’s default settings — minimum memory and the rest as most people start. Change them below for your own case.
Raspberry Pi 5 is the cheapest usable gigabyte at $12.5/GB within your filters. Check its bandwidth before committing — value per gigabyte says what fits, not how fast it runs.
Every catalogued card ranked by cost per usable gigabyte.
Raspberry Pi 5 is the cheapest usable gigabyte at $12.5/GB within your filters. Check its bandwidth before committing — value per gigabyte says what fits, not how fast it runs.
| # | Device | Usable | Price | Per GB |
|---|---|---|---|---|
| 1 | Raspberry Pi 5 · 16 GB | 9.60 GB | $120 | $12.5/GB |
| 2 | RTX 5060 Ti 16GB · 16 GB | 14.7 GB | $429 | $29.1/GB |
| 3 | RTX 4060 Ti 16GB · 16 GB | 14.7 GB | $449 | $30.5/GB |
| 4 | RX 9070 XT · 16 GB | 14.7 GB | $599 | $40.7/GB |
| 5 | RX 7900 XTX · 24 GB | 22.1 GB | $999 | $45.2/GB |
| 6 | RTX 5070 Ti · 16 GB | 14.7 GB | $749 | $50.9/GB |
| 7 | RTX 4070 Ti SUPER · 16 GB | 14.7 GB | $799 | $54.3/GB |
| 8 | MacBook Pro M4 Pro · 64 GB | 48.0 GB | $2,799 | $58.3/GB |
| 9 | MacBook Air M4 · 24 GB | 18.0 GB | $1,199 | $66.6/GB |
| 10 | MacBook Pro M4 Pro · 48 GB | 36.0 GB | $2,399 | $66.6/GB |
| 11 | RTX 5080 · 16 GB | 14.7 GB | $999 | $67.9/GB |
| 12 | RTX 4080 SUPER · 16 GB | 14.7 GB | $999 | $67.9/GB |
| 13 | RTX 3090 · 24 GB | 22.1 GB | $1,499 | $67.9/GB |
| 14 | RTX 5090 · 32 GB | 29.4 GB | $1,999 | $67.9/GB |
Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.
| Setting | Default | What it changes |
|---|---|---|
| Minimum memory | 16 GB | Anywhere from 4 to 192 GB. |
| Budget ceiling | 3000 USD | Anywhere from 100 to 20,000 USD. |
The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.
| Minimum memory | Best value | Price per GB | Usable memory | Bandwidth |
|---|---|---|---|---|
| 4 | Raspberry Pi 5 | $12.5 | 9.60 GB | 17 GB/s |
| 51 | MacBook Pro M4 Pro | $58.3 | 48.0 GB | 273 GB/s |
| 98 | Nothing matches | |||
| 145 | Nothing matches | |||
| 192 | Nothing matches |
Raspberry Pi 5 is the cheapest usable gigabyte at $12.5/GB within your filters. Check its bandwidth before committing — value per gigabyte says what fits, not how fast it runs.
No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.
Dollars per gigabyte is the honest first filter for local AI, because memory decides whether a model runs at all. It deliberately ignores compute: a card that fits the model slowly still beats one that cannot load it.
Usable memory is used rather than nameplate, since drivers, the desktop and fragmentation take a real share — a larger share on unified-memory machines.
These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.
The three situations that bring people to this calculation.
Find the most memory your budget buys.
See where older cards still win.
Get the best capability per pound spent.
Four steps, no account, nothing leaves your browser.
Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.
1 further setting: budget ceiling. Defaults are the common case, so change only what differs for you.
The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.
The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.
The questions people ask about this, answered without hedging.
Usually a previous-generation card with a large frame buffer rather than the newest release. The market prices new architectures at a premium that local inference cannot always use.
No — it deliberately ignores bandwidth, which sets speed. Use it to decide what fits, then check the bandwidth on the shortlist before buying.
Because a model that fits but generates at two tokens per second is unusable for chat. Capacity gets you in the door; bandwidth decides whether you stay.
For memory per slot and for running many cards, sometimes. Per dollar of capability they usually lose badly to consumer cards.
All 50 run on the same arithmetic, so answers across them agree.
Different ways of asking the same question, all resolved above.
Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.