An upgrade buys two different things — capacity and bandwidth — and they do not move together. Only one of them changes what you can run at all.

Computed from this tool’s default settings — current card and the rest as most people start. Change them below for your own case.
The upgrade moves the largest model you can run from 14B to 32B and speeds an 8B model up by about 250%.
What a bigger card actually unlocks, in models and in speed.
The upgrade moves the largest model you can run from 14B to 32B and speeds an 8B model up by about 250%.
Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.
| Setting | Default | What it changes |
|---|---|---|
| Current card | RTX 4060 Ti 16GB · 16 GB | The machine the model runs on. Usable memory decides what fits and memory bandwidth decides how fast it runs, so this moves every figure below. |
| Considering | RTX 4090 · 24 GB | The machine the model runs on. Usable memory decides what fits and memory bandwidth decides how fast it runs, so this moves every figure below. |
| Context length | 8192 tokens | Anywhere from 1,024 to 131,072 tokens. |
The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.
| Current card | Verdict | Memory | Bandwidth | Speed on 8B |
|---|---|---|---|---|
| NVIDIA B200 · 180 GB | No capacity gain | 165.6 GB → 22.1 GB | 7,700 → 1,008 GB/s | 1,369 → 179 tok/s |
| Cerebras WSE-3 · 44 GB on-chip SRAM | No capacity gain | 44.0 GB → 22.1 GB | 21,000,000 → 1,008 GB/s | 3,733,333 → 179 tok/s |
| Raspberry Pi 5 · 16 GB | 8B → 32B | 9.60 GB → 22.1 GB | 17 → 1,008 GB/s | 3 → 179 tok/s |
| RTX 4080 SUPER · 16 GB | 14B → 32B | 14.7 GB → 22.1 GB | 736 → 1,008 GB/s | 131 → 179 tok/s |
| RTX 5070 · 12 GB | 14B → 32B | 11.0 GB → 22.1 GB | 672 → 1,008 GB/s | 119 → 179 tok/s |
Both cards top out at 70B, so this buys speed rather than capability — about -87% on an 8B model. Worth it only if throughput is your constraint.
No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.
An upgrade buys two separate things: capacity, which decides which models fit at all, and bandwidth, which decides how fast they run. They do not move together — a card can have more memory and similar bandwidth, which changes what you can load without changing how it feels.
The largest model that fits is the number worth comparing. Going from a card that tops out at 14B to one that reaches 32B is a category change; going from 32B to 34B is not.
These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.
The three situations that bring people to this calculation.
See whether it changes your ceiling or just your speed.
Judge them on the numbers that matter for inference.
Decide if the delta justifies the cost.
Four steps, no account, nothing leaves your browser.
Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.
2 further settings: considering, context length. Defaults are the common case, so change only what differs for you.
The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.
The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.
The questions people ask about this, answered without hedging.
If it raises the largest model you can run, usually yes — that is a category change. If both cards top out at the same size, you are buying speed, which matters only if throughput is your actual complaint.
VRAM first. A model that does not load cannot be fast, and capacity is the harder constraint to work around.
For capacity they can be, and it is often cheaper per gigabyte. For simplicity one large card wins: no tensor parallelism, no interconnect overhead, no power and cooling headaches.
Somewhat — newer cards get better kernel support and features like flash attention sooner. But capacity and bandwidth still explain most of the difference.
All 50 run on the same arithmetic, so answers across them agree.
Different ways of asking the same question, all resolved above.
Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.