Gaming benchmarks do not predict inference. Two numbers do: how much the card holds, and how fast it reads it. Everything else is close to noise for single-stream decoding.

Computed from this tool’s default settings — first card and the rest as most people start. Change them below for your own case.
RTX 3090 wins on capacity — it runs up to 32B against 14B. Capacity beats speed here: a model that does not load cannot be fast.
Two cards compared on what actually matters: capacity and bandwidth.
RTX 3090 wins on capacity — it runs up to 32B against 14B. Capacity beats speed here: a model that does not load cannot be fast.
| Model | Needs | RTX 3090 | RTX 5070 Ti | Winner |
|---|---|---|---|---|
| 3B | 3.54 GB | 444 tok/s | 425 tok/s | RTX 3090 |
| 8B | 6.89 GB | 166 tok/s | 159 tok/s | RTX 3090 |
| 14B | 10.7 GB | 95 tok/s | 91 tok/s | RTX 3090 |
| 32B | 21.8 GB | 42 tok/s | — | RTX 3090 |
| 70B | 44.5 GB | — | — | Neither |
Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.
| Setting | Default | What it changes |
|---|---|---|
| First card | RTX 3090 · 24 GB | The machine the model runs on. Usable memory decides what fits and memory bandwidth decides how fast it runs, so this moves every figure below. |
| Second card | RTX 5070 Ti · 16 GB | The machine the model runs on. Usable memory decides what fits and memory bandwidth decides how fast it runs, so this moves every figure below. |
| Context length | 8192 tokens | Anywhere from 1,024 to 131,072 tokens. |
The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.
| First card | Better for local AI | NVIDIA B200 | RTX 5070 Ti | Largest model |
|---|---|---|---|---|
| NVIDIA B200 · 180 GB | NVIDIA B200 | 165.6 GB | 14.7 GB | 70B vs 14B |
| Cerebras WSE-3 · 44 GB on-chip SRAM | Cerebras WSE-3 | 44.0 GB | 14.7 GB | 32B vs 14B |
| Raspberry Pi 5 · 16 GB | RTX 5070 Ti | 9.60 GB | 14.7 GB | 8B vs 14B |
| RTX 4080 SUPER · 16 GB | RTX 5070 Ti | 14.7 GB | 14.7 GB | 14B vs 14B |
| RTX 5070 · 12 GB | RTX 5070 Ti | 11.0 GB | 14.7 GB | 14B vs 14B |
NVIDIA B200 wins on capacity — it runs up to 70B against 14B. Capacity beats speed here: a model that does not load cannot be fast.
No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.
Gaming benchmarks do not transfer. Local inference is bound by memory capacity, which decides what loads, and memory bandwidth, which decides how fast it runs. Compute throughput barely enters into single-stream decoding.
This is why an older large-memory card often beats a newer smaller one for this workload, even when it loses every gaming comparison.
These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.
The three situations that bring people to this calculation.
Compare on the metrics that decide inference.
See where a previous generation still wins.
Translate gaming numbers into AI relevance.
Four steps, no account, nothing leaves your browser.
Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.
2 further settings: second card, context length. Defaults are the common case, so change only what differs for you.
The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.
The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.
The questions people ask about this, answered without hedging.
The one with more usable memory, unless both hold the models you want — then the one with more bandwidth. Compute throughput barely enters into it.
Because it has more memory. A previous-generation card with 24GB runs models a newer 12GB card simply cannot load, and capacity is the harder constraint.
Very little for generation, which is memory bound. They matter more for prompt processing and for training, which are compute bound.
Both work through ROCm and Vulkan, and the same capacity-and-bandwidth logic applies. The trade-off is software maturity rather than the hardware arithmetic.
All 50 run on the same arithmetic, so answers across them agree.
Different ways of asking the same question, all resolved above.
Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.