All tools
Cost · free, no sign-up

How much would running locally save me?

Local inference is often described as free because the electricity is cheap. The hardware is not, and whether local actually saves money depends almost entirely on how hard you work the card.

7 inputs4 questions answeredUpdated for 2026 hardware
Local LLM Cost Savings Calculator — How much would running locally save me?
Answer first

The short answer

Computed from this tool’s default settings — input tokens per month and the rest as most people start. Change them below for your own case.

Monthly saving$1,497

Running this volume locally costs about $3.20 of electricity against $1,500 on the API, so the card pays for itself in about 1.1 months.

The calculator

Local LLM Cost Savings Calculator

Compare an API bill against electricity, and find the break-even point.

Your setup

Frontier models are around $10. Mid-tier is nearer $1.40.

Actual inference time, not hours the machine is on.

Monthly saving$1,497

Running this volume locally costs about $3.20 of electricity against $1,500 on the API, so the card pays for itself in about 1.1 months.

API bill$1,500
Electricity$3.20320 W for 40 h
Hardware$1,599
Pays for itself in1.1 months
Inputs

What each setting changes

Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.

SettingDefaultWhat it changes
Input tokens per month100 millionsYour own figure in millions, starting from 100. Change it to match what you actually run.
Output tokens per month10 millionsYour own figure in millions, starting from 10. Change it to match what you actually run.
API input price10 $ / M tokensFrontier models are around $10. Mid-tier is nearer $1.40.
API output price50 $ / M tokensYour own figure in $ / M tokens, starting from 50. Change it to match what you actually run.
Your hardwareRTX 4090 · 24 GBThe machine the model runs on. Usable memory decides what fits and memory bandwidth decides how fast it runs, so this moves every figure below.
Electricity price0.25 $ / kWhYour own figure in $ / kWh, starting from 0.25. Change it to match what you actually run.
Hours of generation per month40 hoursActual inference time, not hours the machine is on.
Worked examples

Real answers across your hardware

The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.

Your hardwareMonthly savingAPI billElectricity
NVIDIA B200 · 180 GB$1,497$1,500$3.20
Cerebras WSE-3 · 44 GB on-chip SRAM$1,497$1,500$3.20
Raspberry Pi 5 · 16 GB$1,497$1,500$3.20$120
RTX 4080 SUPER · 16 GB$1,497$1,500$3.20$999
RTX 5070 · 12 GB$1,497$1,500$3.20$549

Running this volume locally costs about $3.20 of electricity against $1,500 on the API.

Method

How this is calculated

No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.

The hosted side is simply tokens times price. The local side is electricity: the card’s power draw times the hours it actually generates, times your tariff. Hardware is treated as a separate up-front cost, shown as a payback period rather than folded into the monthly figure.

This compares cost only. A local model will not match a frontier model on the hardest work — the honest use is to move the volume that does not need one.

These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.

Use cases

Who this is for

The three situations that bring people to this calculation.

Use case 01

Justifying a build

See whether the numbers support the purchase.

Use case 02

Comparing to an API

Put both on the same per-token basis.

Use case 03

Heavy users

Find the volume where local clearly wins.

Walkthrough

How to use this calculator

Four steps, no account, nothing leaves your browser.

  1. Set input tokens per month

    Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.

  2. Adjust the rest to match your setup

    6 further settings: output tokens per month, api input price, api output price, your hardware, electricity price, hours of generation per month. Defaults are the common case, so change only what differs for you.

  3. Read the headline, then the table

    The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.

  4. Check it against the method

    The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.

Questions

How much would running locally save me: common questions

The questions people ask about this, answered without hedging.

Is local AI cheaper than an API?

Per token, only at high utilisation. Electricity alone is a fraction of API pricing, but adding hardware amortisation flips it for light use. A card idle 95% of the day is expensive per token however efficient it is.

What is the break-even point?

It depends on the card price and the API rate you are avoiding, but for a consumer GPU against cheap hosted models it is usually millions of tokens a month sustained.

What does local buy that an API does not?

Privacy, offline capability, no rate limits, no per-token anxiety while experimenting, and a model that cannot be deprecated out from under you. These are usually the real reasons, and they do not appear in a cost calculation.

Should I include my time?

If you are being honest, yes. Local setup and maintenance cost hours that an API call does not, which matters more than electricity at most volumes.

The rest of the set

All 50 run on the same arithmetic, so answers across them agree.

Coverage

Searches this page answers

Different ways of asking the same question, all resolved above.

local llm vs api costis running llm locally cheaperllm cost comparison calculatorself host llm costgpu vs openai api costlocal ai cost savingscost per token localhow much would running locally save melocal llm cost savings calculatorlocal llm cost savings calculator onlinefree local llm cost savings calculator
Next step

Now find the models that fit

Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.