The case for buying a GPU is usually made with a payback period, and the number is almost always worse than expected because the card sits idle most of the day.

Computed from this tool’s default settings — your hardware and the rest as most people start. Change them below for your own case.
At 500k tokens a day the API costs about $9.12 a month. A $1,599 card never recovers that on token cost alone — buy it for privacy, latency or offline use instead.
Break-even in months, from your real token volume.
At 500k tokens a day the API costs about $9.12 a month. A $1,599 card never recovers that on token cost alone — buy it for privacy, latency or offline use instead.
Every input moves the result for a reason. This is what each one does and where to find the value for your own machine.
| Setting | Default | What it changes |
|---|---|---|
| Your hardware | RTX 4090 · 24 GB | The machine the model runs on. Usable memory decides what fits and memory bandwidth decides how fast it runs, so this moves every figure below. |
| Model size | 8B | 6 options, from 3B to 405B. |
| Power draw under load | 350 W | Board power while generating, not idle. Card TDP is a good starting point. |
| Electricity price | 0.18 per kWh | Your actual tariff. US average is near $0.17, UK near £0.25, India near ₹8. |
| API price you would pay | 0.6 per 1M tokens | Blend input and output. Small hosted models sit near $0.20–$1.00. |
| Your usage | 500 thousand tokens/day | Your own figure in thousand tokens/day, starting from 500. Change it to match what you actually run. |
The same calculation run at a range of settings, with everything else left at its default. These are computed by the tool itself, not written by hand.
| Model size | Payback period | Card price | API per month | Power per month |
|---|---|---|---|---|
| 3B | Never at this volume | $1,599 | $9.12 | $0.56 |
| 8B | Never at this volume | $1,599 | $9.12 | $1.48 |
| 14B | Never at this volume | $1,599 | $9.12 | $2.60 |
| 32B | Never at this volume | $1,599 | $9.12 | $5.94 |
| 70B | Never at this volume | $1,599 | $9.12 | $12.99 |
At 500k tokens a day the API costs about $9.12 a month. A $1,599 card never recovers that on token cost alone — buy it for privacy, latency or offline use instead.
No lookup tables and no invented constants. Here is the arithmetic, so you can check it against your own numbers.
Break-even is the purchase price divided by the monthly saving, where the saving is what the API would have charged minus what the electricity costs. If electricity exceeds the API price there is no break-even, and the tool says so rather than printing a large number.
API prices are inputs rather than constants, because providers change them and a hard-coded number would quietly go wrong. The defaults were current when this tool was written — check your provider’s pricing page and correct them if they have moved.
These are well-founded engineering estimates, not benchmark results. Your quantisation, runtime and context length all move the real number, and usable memory is an assumption rather than a specification. See the full methodology for every assumption behind these figures.
The three situations that bring people to this calculation.
See the break-even at your real token volume.
Confirm the volume justifies it.
Find out the payback never arrives, before spending.
Four steps, no account, nothing leaves your browser.
Start at the top of the panel. Every figure recalculates as you change it — there is no submit button, because watching the number move is the point.
5 further settings: model size, power draw under load, electricity price, api price you would pay, your usage. Defaults are the common case, so change only what differs for you.
The large figure answers the question. The table underneath shows how the answer changes across nearby settings, which is usually where the decision actually gets made.
The arithmetic is written out above. If a number looks wrong for your hardware, the assumptions are the first place to look — usable memory and quantisation are the two that vary most.
The questions people ask about this, answered without hedging.
It depends entirely on volume. At a few hundred thousand tokens a day against cheap hosted pricing, often never. At millions a day against premium pricing, months.
Because at your volume the electricity cost approaches or exceeds what the API would charge. When that happens there is no saving to recover the purchase from, and saying so is more useful than printing a large number.
Possibly. Privacy, offline use, no rate limits and unmetered experimentation are real benefits that a token-cost comparison cannot price. Just buy it for those reasons rather than a spreadsheet that does not hold up.
It should. GPUs hold value better than most hardware, so the true cost is the purchase price minus what you will recover — which can materially shorten the payback.
All 50 run on the same arithmetic, so answers across them agree.
Different ways of asking the same question, all resolved above.
Sizing is only half the problem. Model Radar takes your hardware and shows which models actually run on it, ranked by what they are good at — the same arithmetic as this page, applied to every model worth running.