The real choice is not portable versus not — it is unified memory versus discrete. They fail in opposite directions, and which one suits you depends on model size.

Computed from this tool’s default settings — model size and the rest as most people start. Change them below for your own case.
Within $2,500, a discrete card runs a 8B model faster than anything portable. Buy the desktop unless you genuinely need to carry it.
Capacity, speed and cost across both, for the model you want to run.
Within $2,500, a discrete card runs a 8B model faster than anything portable. Buy the desktop unless you genuinely need to carry it.
| Form | Device | Usable | Speed | Price |
|---|---|---|---|---|
| Unified (portable) | MacBook Pro M4 Pro | 36.0 GB | 49 tok/s | $2,399 |
| Discrete (desktop) | RTX 5090 | 29.4 GB | 319 tok/s | $1,999 |
Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.
| Setting | Default | What it changes |
|---|---|---|
| Model size | 8B | 6 options, from 3B to 405B. |
| Context length | 8192 tokens | Anywhere from 1,024 to 131,072 tokens. |
| Budget | 2500 USD | Anywhere from 500 to 15,000 USD. |
The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.
| Model size | Faster within budget | Model needs | Options in budget | Best discrete |
|---|---|---|---|---|
| 3B | Desktop | 3.54 GB | 19 | 850 tok/s |
| 8B | Desktop | 6.89 GB | 17 | 319 tok/s |
| 14B | Desktop | 10.7 GB | 16 | 182 tok/s |
| 32B | Desktop | 21.8 GB | 5 | 80 tok/s |
| 70B | Nothing fits this budget | 44.5 GB | 0 |
Within $2,500, a discrete card runs a 3B model faster than anything portable. Buy the desktop unless you genuinely need to carry it.
No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.
The comparison that matters is not the badge but the memory architecture. Unified-memory laptops trade bandwidth for capacity — they hold large models slowly. Discrete desktop cards do the reverse: far more bandwidth per gigabyte, with a hard ceiling on how much fits.
Mobile discrete GPUs also carry less memory than their desktop namesakes, so match on the memory figure rather than the model number.
These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.
The three situations that bring people to this calculation.
Decide where the money goes.
See what mobility actually costs in speed.
Find which form factor holds them at all.
Four steps, no account, nothing leaves your browser.
Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.
2 further settings: context length, budget. Defaults are the common case, so change only what differs for you.
The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.
The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.
The questions people ask about this, answered without hedging.
A unified-memory laptop with plenty of RAM holds large models well, just slowly. A gaming laptop with a discrete GPU is fast but usually memory-limited, since mobile cards carry less VRAM than their desktop namesakes.
Mobile parts run at lower power and often lower memory bandwidth, and frequently ship with less VRAM despite the same model number. Match on memory and bandwidth, not name.
A desktop gives more capability per pound, every time. Buy a laptop when you genuinely need to carry it, not because the specs look comparable.
Workable but compromised: the enclosure and cable limit bandwidth to the card, and support varies by platform. Better than nothing, well short of a desktop slot.
All 50 run on the same arithmetic, so answers across them agree.
Different ways of asking the same question, all resolved above.
Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.