APIs quote a price per million tokens, so that is the only basis on which local can honestly be compared. Doing it properly means counting the card, not just the electricity.

Computed from this tool’s default settings — your hardware and the rest as most people start. Change them below for your own case.
At 10% utilisation, RTX 4090 works out at $0.81 per million tokens all-in — of which only $0.098 is electricity. Amortisation dominates until the card is genuinely busy.
Electricity plus hardware amortisation, expressed the way APIs price.
At 10% utilisation, RTX 4090 works out at $0.81 per million tokens all-in — of which only $0.098 is electricity. Amortisation dominates until the card is genuinely busy.
| Utilisation | Tokens/year | Cost per 1M |
|---|---|---|
| 5% | 283M | $1.51 |
| 10% | 565M | $0.81 |
| 25% | 1,413M | $0.38 |
| 50% | 2,826M | $0.24 |
| 100% | 5,651M | $0.17 |
Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.
| Setting | Default | What it changes |
|---|---|---|
| Your hardware | RTX 4090 · 24 GB | The machine the model runs on. Usable memory decides what fits and memory bandwidth decides how fast it runs, so this moves every figure below. |
| Model size | 8B | 6 options, from 3B to 405B. |
| Power draw under load | 350 W | Board power while generating, not idle. Card TDP is a good starting point. |
| Electricity price | 0.18 per kWh | Your actual tariff. US average is near $0.17, UK near £0.25, India near ₹8. |
| Hardware lifetime | 4 years | Anywhere from 1 to 10 years. |
| Utilisation | 10 % of the time generating | A card that idles 90% of the day still costs its full purchase price. |
The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.
| Model size | All-in cost per 1M tokens | Electricity only | Throughput | Tokens per year |
|---|---|---|---|---|
| 3B | $0.30 | $0.037 | 478 tok/s | 1,507M |
| 8B | $0.81 | $0.098 | 179 tok/s | 565M |
| 14B | $1.41 | $0.171 | 102 tok/s | 323M |
| 32B | $3.22 | $0.391 | 45 tok/s | 141M |
| 70B | $7.04 | $0.854 | 20 tok/s | 65M |
At 10% utilisation, RTX 4090 works out at $0.30 per million tokens all-in — of which only $0.037 is electricity. Amortisation dominates until the card is genuinely busy.
No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.
Local inference has two costs: electricity while generating, and the hardware itself spread over its useful life. Comparing only electricity to an API price is the mistake that makes local look free.
Utilisation is what usually decides the answer. Amortised hardware cost per token falls as the card works harder, so a GPU generating a few thousand tokens a day is expensive per token no matter how efficient it is.
These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.
The three situations that bring people to this calculation.
Put local on the same units as a provider price list.
See what utilisation the case actually depends on.
Work out your own cost basis before pricing anything.
Four steps, no account, nothing leaves your browser.
Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.
5 further settings: model size, power draw under load, electricity price, hardware lifetime, utilisation. Defaults are the common case, so change only what differs for you.
The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.
The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.
The questions people ask about this, answered without hedging.
Per token, only at high utilisation. Electricity alone is usually a fraction of API pricing, but adding hardware amortisation flips it for light use. Local wins decisively on privacy, offline capability and unmetered experimentation.
Electricity alone is usually cents per million tokens. Including the card spread over its life, the all-in figure is dominated by amortisation until the GPU is genuinely busy — which for most people it is not.
The card costs the same whether it runs all day or five minutes. Spread over few tokens, that fixed cost is enormous per token; spread over many, it disappears.
If the machine exists anyway, counting the GPU is fair. If you bought it for this, count all of it — including the power supply and cooling the card required.
At sustained high volume against premium model pricing, yes. Against the cheapest hosted small models, rarely — those are priced close to the metal.
All 50 run on the same arithmetic, so answers across them agree.
Different ways of asking the same question, all resolved above.
Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.