Contents
Tags

Hardware advice usually goes model-first: pick the model, buy what runs it. That works until you pick a size that does not land cleanly on any tier — and two popular sizes do exactly that.
At Q4_K_M with an 8K context, against usable memory:
49B and 70B are the dead zone. Each one lands just past a tier boundary, so the cheapest hardware that runs it is a whole category up — and the category up is roughly double the money.
70B is the most recommended 'serious' local model, and it is the worst fit in the range. Two 24 GB cards get you 43.2 GB usable against a 44.5 GB requirement — short by 1.3 GB, which means a lower quantisation, a shorter context, or a third card.
The alternatives are all a step change in cost: a 48 GB workstation card, three consumer cards with the tensor-parallel awkwardness that brings, or unified memory at 64 GB and above where capacity is cheap and bandwidth is the compromise.
Check the requirement against the tier boundary before you fall in love with a model size. Landing 1 GB over a boundary costs roughly twice as much as landing 1 GB under it, and the difference in what the model can do is nothing like a factor of two.
Every figure here is computed the same way the calculators do it: weights are parameters times bits-per-weight divided by eight, plus a KV-cache term for your context, plus about a gigabyte of runtime overhead. Usable memory is roughly 90% of nameplate on a discrete card.
Find the cheapest hardware that actually holds the model you want.
Start from the model →Tools
Find AI models that fit your exact hardware. Enter your specs and get a ranked list instantly.
Newsletter