All tools
Cost · free, no sign-up

What does it actually cost to run a model locally?

APIs quote a price per million tokens, so that is the only basis on which local can honestly be compared. Doing it properly means counting the card, not just the electricity.

6 inputs5 questions answeredUpdated for 2026 hardware
Local Cost per Million Tokens — What does it actually cost to run a model locally?
Answer first

The short answer

Computed from this tool’s default settings — your hardware and the rest as most people start. Change them below for your own case.

All-in cost per 1M tokens$0.81

At 10% utilisation, RTX 4090 works out at $0.81 per million tokens all-in — of which only $0.098 is electricity. Amortisation dominates until the card is genuinely busy.

The calculator

Local Cost per Million Tokens

Electricity plus hardware amortisation, expressed the way APIs price.

Your setup

Board power while generating, not idle. Card TDP is a good starting point.

Your actual tariff. US average is near $0.17, UK near £0.25, India near ₹8.

10

A card that idles 90% of the day still costs its full purchase price.

All-in cost per 1M tokens$0.81

At 10% utilisation, RTX 4090 works out at $0.81 per million tokens all-in — of which only $0.098 is electricity. Amortisation dominates until the card is genuinely busy.

Electricity only$0.098What people usually quote
Throughput179 tok/s
Tokens per year565M
Hardware per year$400
UtilisationTokens/yearCost per 1M
5%283M$1.51
10%565M$0.81
25%1,413M$0.38
50%2,826M$0.24
100%5,651M$0.17
Inputs

What each setting changes

Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.

SettingDefaultWhat it changes
Your hardwareRTX 4090 · 24 GBThe machine the model runs on. Usable memory decides what fits and memory bandwidth decides how fast it runs, so this moves every figure below.
Model size8B6 options, from 3B to 405B.
Power draw under load350 WBoard power while generating, not idle. Card TDP is a good starting point.
Electricity price0.18 per kWhYour actual tariff. US average is near $0.17, UK near £0.25, India near ₹8.
Hardware lifetime4 yearsAnywhere from 1 to 10 years.
Utilisation10 % of the time generatingA card that idles 90% of the day still costs its full purchase price.
Worked examples

Real answers across model size

The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.

Model sizeAll-in cost per 1M tokensElectricity onlyThroughputTokens per year
3B$0.30$0.037478 tok/s1,507M
8B$0.81$0.098179 tok/s565M
14B$1.41$0.171102 tok/s323M
32B$3.22$0.39145 tok/s141M
70B$7.04$0.85420 tok/s65M

At 10% utilisation, RTX 4090 works out at $0.30 per million tokens all-in — of which only $0.037 is electricity. Amortisation dominates until the card is genuinely busy.

Method

How this is calculated

No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.

Local inference has two costs: electricity while generating, and the hardware itself spread over its useful life. Comparing only electricity to an API price is the mistake that makes local look free.

Utilisation is what usually decides the answer. Amortised hardware cost per token falls as the card works harder, so a GPU generating a few thousand tokens a day is expensive per token no matter how efficient it is.

These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.

Use cases

Who this is for

The three situations that bring people to this calculation.

Use case 01

Comparing to an API

Put local on the same units as a provider price list.

Use case 02

Justifying hardware

See what utilisation the case actually depends on.

Use case 03

Running a service

Work out your own cost basis before pricing anything.

Walkthrough

How to use this calculator

Four steps, no account, nothing leaves your browser.

  1. Set your hardware

    Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.

  2. Adjust the rest to match your setup

    5 further settings: model size, power draw under load, electricity price, hardware lifetime, utilisation. Defaults are the common case, so change only what differs for you.

  3. Read the headline, then the table

    The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.

  4. Check it against the method

    The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.

Questions

What does it actually cost to run a model locally: common questions

The questions people ask about this, answered without hedging.

Is running an LLM locally cheaper than an API?

Per token, only at high utilisation. Electricity alone is usually a fraction of API pricing, but adding hardware amortisation flips it for light use. Local wins decisively on privacy, offline capability and unmetered experimentation.

What does it cost to run an LLM locally?

Electricity alone is usually cents per million tokens. Including the card spread over its life, the all-in figure is dominated by amortisation until the GPU is genuinely busy — which for most people it is not.

Why does utilisation matter so much?

The card costs the same whether it runs all day or five minutes. Spread over few tokens, that fixed cost is enormous per token; spread over many, it disappears.

Should I count the whole PC or just the GPU?

If the machine exists anyway, counting the GPU is fair. If you bought it for this, count all of it — including the power supply and cooling the card required.

Does the comparison ever favour local on cost alone?

At sustained high volume against premium model pricing, yes. Against the cheapest hosted small models, rarely — those are priced close to the metal.

The rest of the set

All 50 run on the same arithmetic, so answers across them agree.

Coverage

Searches this page answers

Different ways of asking the same question, all resolved above.

local llm cost per million tokenscost to run llm locallyself hosted llm cost per tokengpu cost per tokenlocal inference cost calculatorllm electricity cost per tokenamortised gpu cost aiwhat does it actually cost to run a model locallylocal cost per million tokenslocal cost per million tokens onlinefree local cost per million tokens
Next step

Now find the models that fit

Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.