RUNYARD.DEV / COMPARE
Pick two devices · Device B can be wrapped with TurboQuant
Two numbers decide whether a local model is usable on a given machine, and this page compares both. Memory capacity decides whether the model loads at all. Memory bandwidth decides how fast it generates once it has, because producing a single token means reading every active weight once.
Compute throughput barely enters into single-stream decoding, which is why two cards with similar TFLOPs can feel completely different, and why an older card with more memory often beats a newer one with less.
Every figure on this page comes from the same arithmetic used across the rest of the site, so a verdict here agrees with the calculators and the hardware pages.
These are engineering estimates, not benchmark results. Quantisation, runtime and context length all move the real number. See the methodology for every assumption behind them.
Because it has more memory. A previous-generation card with 24 GB runs models a newer 12 GB card cannot load at all, and a model that does not fit does not run slowly — it does not run. Capacity is the harder constraint, so it is the first thing to compare.
Capacity first, bandwidth second. Work out which models each device can hold, and only then compare speed on the models both can run. Comparing tokens per second between two devices that fit different models is comparing different products.
Apple Silicon shares one pool between CPU, GPU and the system, so capacity is cheap and bandwidth is what separates the chip tiers. A high-end unified machine holds far larger models than a consumer card at similar money, and runs them more slowly. Which one wins depends entirely on the size of model you need.
It applies a quantisation profile to device B only, so you can see what dropping precision buys in fit and context on one side of the comparison while the other stays fixed. Lower precision means smaller weights, which means both more headroom and faster generation — at some cost in quality.
It lets a larger model load, but it does not make it fast. Layers that spill to system RAM run at a fraction of card bandwidth, and because every token passes through every layer, the slow side dominates. A lower quantisation that fits entirely on the card is usually faster than offloading.
If you already know which model you want, it is usually quicker to start from the model and let the arithmetic name the hardware: