Contents
Tags

Almost every version of this comparison you will read is rigged, usually without meaning to be. It takes the electricity cost of running a model at home, compares it to the list price of a frontier API, and concludes that local is a hundred times cheaper. Both halves of that are wrong: the local figure leaves out the hardware, and the two sides are not the same product.
Here is the comparison done properly, with the arithmetic shown.
A realistic self-host: a used RTX 3090 at $900, running a 14B model at Q4_K_M. That is 24 GB at 936 GB/s, which puts it at about 95 tokens per second on that model. Electricity at $0.18 per kWh, 350W under load, and a four-year life on the card.
Local has two costs, and the second is the one that gets dropped:
At 10% utilisation, meaning the GPU is actually generating for about two and a half hours a day, that card produces roughly 300 million tokens a year and the all-in cost is about $0.93 per million tokens. Five times the electricity-only figure. At 2% utilisation it is $3.94. At 50% it falls to $0.33.
Utilisation is the whole game. The card costs the same whether it runs all day or five minutes, so the fixed cost per token is enormous when spread thin and disappears when the machine is genuinely busy.
This is where the comparison usually cheats, because API prices span two orders of magnitude depending on which model you pick:
Only against the expensive end, and only at volume. Break-even on that $900 card, after subtracting the electricity you still pay:
Read the first line again, because it is the one that matters most and the one nobody prints. If the thing you would otherwise use is a cheap hosted small model, buying hardware to replace it will not pay for itself in the life of the card. Those models are already priced near the cost of serving them.
The frontier row above is the one that makes local look good, and it is also the one that is not a like-for-like swap. A 14B model running on your desk is not a substitute for a frontier model. It is a different, weaker product that happens to be cheaper.
If you replace frontier API calls with a local 14B and your task still works, the honest conclusion is that you were overpaying for capability you did not need — and the cheapest fix was probably a cheaper hosted model, not a GPU. If the task stops working, you have not saved anything at all.
The fair comparison is a local model against a hosted model of similar capability. On that basis the API wins on cost almost every time, because the provider is amortising hardware across thousands of users and you are amortising it across one.
Because cost is not why local wins, and treating it as the argument is what makes the argument lose.
Those are real and they are worth money. They are just not the same argument as dollars per million tokens, and mixing the two produces the breathless posts that get this wrong.
Put your own utilisation, tariff and card price in and see the all-in cost per million tokens.
Work out your local cost →Or start from the API bill and find the volume where a card pays back.
Open the payback calculator →Tools
Newsletter