All tools
Hardware · free, no sign-up

What is the cheapest GPU that runs this model?

Work backwards. Pick the model you want to run, and the arithmetic names the cheapest card that holds it — often a rung lower than people assume.

2 inputs4 questions answeredUpdated for 2026 hardware
Cheapest GPU for a Model — What is the cheapest GPU that runs this model?
Answer first

The short answer

Computed from this tool’s default settings — model size and the rest as most people start. Change them below for your own case.

Cheapest at Q4_K_MRaspberry Pi 5

A 8B model at Q4_K_M needs 6.89 GB, which Raspberry Pi 5 covers for $120 at roughly 3 tok/s.

The calculator

Cheapest GPU for a Model

The lowest-cost card that fits, at every quantisation level.

Your setup
8,192
Cheapest at Q4_K_MRaspberry Pi 5

A 8B model at Q4_K_M needs 6.89 GB, which Raspberry Pi 5 covers for $120 at roughly 3 tok/s.

Memory needed6.89 GB
Price$120
Usable memory9.60 GB
Expected speed3 tok/s
QuantisationNeedsCheapest cardPrice
Q8_010.9 GBArc B580$249
Q6_K8.99 GBRaspberry Pi 5$120
Q5_K_M8.09 GBRaspberry Pi 5$120
Q4_K_M6.89 GBRaspberry Pi 5$120
Q3_K_M5.89 GBRaspberry Pi 5$120
Inputs

What each setting changes

Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.

SettingDefaultWhat it changes
Model size8B6 options, from 3B to 405B.
Context length8192 tokensAnywhere from 1,024 to 131,072 tokens.
Worked examples

Real answers across model size

The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.

Model sizeCheapest at Q4_K_MMemory neededPriceUsable memory
3BRaspberry Pi 53.54 GB$804.80 GB
8BRaspberry Pi 56.89 GB$1209.60 GB
14BArc B58010.7 GB$24911.0 GB
32BRX 7900 XTX21.8 GB$99922.1 GB
70BMacBook Pro M4 Pro44.5 GB$2,79948.0 GB

A 3B model at Q4_K_M needs 3.54 GB, which Raspberry Pi 5 covers for $80 at roughly 8 tok/s.

Method

How this is calculated

No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.

For each quantisation the tool computes total memory — weights, KV cache and runtime overhead — then picks the cheapest catalogued device whose usable memory covers it.

Dropping one rung down the quantisation ladder often changes which card you need by a whole price tier, which is why the table shows every rung rather than a single recommendation.

These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.

Use cases

Who this is for

The three situations that bring people to this calculation.

Use case 01

Target a model

Find the minimum card for a specific model.

Use case 02

Trading quality

See how much a lower quantisation saves in hardware.

Use case 03

Shopping

Turn a model choice into a purchase decision.

Walkthrough

How to use this calculator

Four steps, no account, nothing leaves your browser.

  1. Set model size

    Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.

  2. Adjust the rest to match your setup

    1 further setting: context length. Defaults are the common case, so change only what differs for you.

  3. Read the headline, then the table

    The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.

  4. Check it against the method

    The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.

Questions

What is the cheapest GPU that runs this model: common questions

The questions people ask about this, answered without hedging.

What is the cheapest GPU that runs a 70B model?

At Q4_K_M a 70B needs roughly 40 GB plus cache, so a 48GB card or two 24GB cards. Below that you are into partial offload, which is slow enough to change the answer.

How much does quantisation change the card I need?

Often by a whole price tier. Moving from Q8 to Q4 nearly halves the requirement, which is regularly the difference between one card and two.

Should I buy for today’s models or tomorrow’s?

Buy memory. Model sizes have not shrunk, and extra capacity is what keeps a card useful as new releases arrive.

Is a used card a reasonable answer here?

Often the best one. Previous-generation cards with large frame buffers are the value sweet spot for inference — just test the memory under load before paying.

The rest of the set

All 50 run on the same arithmetic, so answers across them agree.

Coverage

Searches this page answers

Different ways of asking the same question, all resolved above.

cheapest gpu for 70bminimum gpu for llamawhat gpu do i need for 32b modelbudget gpu for local llmcheapest way to run 70bgpu requirements by model sizeminimum vram for modelwhat is the cheapest gpu that runs this modelcheapest gpu for a modelcheapest gpu for a model onlinefree cheapest gpu for a model
Next step

Now find the models that fit

Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.