All tools
Will it fit · free, no sign-up

What is the best model my hardware can run?

The best model is not the biggest one that loads — it is the biggest one that loads at a quantisation worth using and still generates faster than you read.

2 inputs4 questions answeredUpdated for 2026 hardware
Best Model for My Hardware — What is the best model my hardware can run?
Answer first

The short answer

Computed from this tool’s default settings — your hardware and the rest as most people start. Change them below for your own case.

Largest worth running32B at Q4_K_M

RTX 4090 runs up to a 32B model at Q4_K_M. Anything larger only fits at a quantisation that degrades noticeably.

The calculator

Best Model for My Hardware

The largest model that fits at a quantisation worth running.

Your setup
8,192

The KV cache grows linearly with this. It is the biggest lever you have.

Largest worth running32B at Q4_K_M

RTX 4090 runs up to a 32B model at Q4_K_M. Anything larger only fits at a quantisation that degrades noticeably.

Usable memory22.1 GB
Bandwidth1,008 GB/s
Expected speed45 tok/s
Model sizeBest quantMemorySpeed
1BQ8_02.55 GB759 tok/s
3BQ8_05.04 GB253 tok/s
7BQ8_09.74 GB108 tok/s
8BQ8_010.9 GB95 tok/s
14BQ8_017.7 GB54 tok/s
27BQ4_K_M18.7 GB53 tok/s
32BQ4_K_M21.8 GB45 tok/s
70BDoes not fit
120BDoes not fit
405BDoes not fit
Inputs

What each setting changes

Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.

SettingDefaultWhat it changes
Your hardwareRTX 4090 · 24 GBThe machine the model runs on. Usable memory decides what fits and memory bandwidth decides how fast it runs, so this moves every figure below.
Context length8192 tokensThe KV cache grows linearly with this. It is the biggest lever you have.
Worked examples

Real answers across your hardware

The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.

Your hardwareLargest worth runningUsable memoryBandwidthExpected speed
NVIDIA B200 · 180 GB120B at Q8_0165.6 GB7,700 GB/s91 tok/s
Cerebras WSE-3 · 44 GB on-chip SRAM32B at Q8_044.0 GB21,000,000 GB/s933,333 tok/s
Raspberry Pi 5 · 16 GB8B at Q6_K9.60 GB17 GB/s3 tok/s
RTX 4080 SUPER · 16 GB14B at Q6_K14.7 GB736 GB/s75 tok/s
RTX 5070 · 12 GB14B at Q4_K_M11.0 GB672 GB/s68 tok/s

NVIDIA B200 runs up to a 120B model at Q8_0. Anything larger only fits at a quantisation that degrades noticeably.

Method

How this is calculated

No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.

Memory for a model is three things added together: the weights, which are parameters × bits-per-weight ÷ 8; the KV cache, which grows linearly with context length; and about a gigabyte of runtime overhead for the CUDA context, activations and framework.

Capacity follows a model’s total parameter count even for mixture-of-experts designs, because the router may select any expert on the next token and all of them must stay resident. Only throughput follows the active count.

Models that only fit at Q3_K_M are excluded from the recommendation. The degradation there is visible, and a smaller model at Q4_K_M is usually the better choice than a larger one at Q3.

These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.

Use cases

Who this is for

The three situations that bring people to this calculation.

Use case 01

Getting started

A shortlist for your specific card.

Use case 02

After an upgrade

Find what the new hardware unlocks.

Use case 03

Picking a daily driver

Balance capability against responsiveness.

Walkthrough

How to use this calculator

Four steps, no account, nothing leaves your browser.

  1. Set your hardware

    Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.

  2. Adjust the rest to match your setup

    1 further setting: context length. Defaults are the common case, so change only what differs for you.

  3. Read the headline, then the table

    The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.

  4. Check it against the method

    The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.

Questions

What is the best model my hardware can run: common questions

The questions people ask about this, answered without hedging.

What is the best local LLM for my GPU?

Whatever fits at Q4_K_M or better with room for your context, and still generates at reading speed. On 8GB that is a good 7–8B model; on 24GB a 32B; on 48GB or more a 70B.

Should I run the biggest model that fits?

Only if it fits at Q4 or higher. A 32B squeezed into Q3 is usually worse than a 14B at Q5, and it will be slower as well.

Does the best model depend on the task?

Yes. Coding models beat general ones at code even when smaller, and a 7B tuned for your domain can beat a 70B generalist. Size is the constraint, not the goal.

How often does the answer change?

Frequently — a new release can move the ceiling for a given card. Check again after any major model launch rather than assuming last year’s pick still holds.

The rest of the set

All 50 run on the same arithmetic, so answers across them agree.

Coverage

Searches this page answers

Different ways of asking the same question, all resolved above.

best llm for my gpubest local model for 24gb vramwhat model should i runbest llm for 8gb vrambest local ai model 2026recommended llm by gpubest model for rtx 4090what is the best model my hardware can runbest model for my hardwarebest model for my hardware onlinefree best model for my hardware
Next step

Now find the models that fit

Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.