The RTX 5070 Ti pairs 16 GB of VRAM with 896 GB/s of memory bandwidth. Those two numbers, in that order, determine everything about what it can run and how quickly.
Of the 53 open-weight models we track, 38 fit entirely in this card's memory at some usable quantisation. The largest is Qwen3 30B A3B Instruct 2507 at 30B parameters, which works because only 3B are active per token.
Highest-fidelity quantisation that fits in 16 GB, with an 8K context window:
| Model | Parameters | Best quant | VRAM used | Est. speed |
|---|---|---|---|---|
| Qwen3 30B A3B Instruct 2507 | 30B / 3B active | Q3_K_M | 15.0 GB | ~496 tok/s |
| Nemotron 3.5 Lightning | 30B / 3B active | Q3_K_M | 15.0 GB | ~546 tok/s |
| Cohere North Mini Code | 30B / 3B active | Q3_K_M | 15.0 GB | ~546 tok/s |
| Qwen3.8 27B | 27B | Q3_K_M | 15.4 GB | ~61 tok/s |
| Qwen3.6 27B | 27B | Q3_K_M | 15.4 GB | ~61 tok/s |
| Qwen3.5-27B | 27B | Q3_K_M | 15.4 GB | ~61 tok/s |
| Gemma 2 27B | 27B | Q3_K_M | 15.4 GB | ~61 tok/s |
| Gemma 4 26B A4B | 25B | Q3_K_M | 14.5 GB | ~65 tok/s |
| Mistral Small 3.1 24B | 24B | Q3_K_M | 13.9 GB | ~68 tok/s |
| Voxtral Small 24B 2507 | 24B | Q3_K_M | 13.9 GB | ~68 tok/s |
| Mistral Small 3.2 24B | 24B | Q3_K_M | 13.9 GB | ~68 tok/s |
| Mistral Small 3 | 24B | Q3_K_M | 13.9 GB | ~68 tok/s |
| gpt-oss-20b | 21B / 4B active | Q4_K_M | 13.7 GB | ~354 tok/s |
| Llama 4 Scout | 17B | Q5_K_M | 15.1 GB | ~59 tok/s |
| Llama 4 Maverick | 17B | Q5_K_M | 15.1 GB | ~59 tok/s |
| Qwen3 14B | 14B | Q6_K | 14.4 GB | ~62 tok/s |
| Ministral 3 14B 2512 | 14B | Q6_K | 14.4 GB | ~62 tok/s |
| Hunyuan A13B Instruct | 13B / 13B active | Q6_K | 13.5 GB | ~67 tok/s |
| Mistral Mistral Nemo | 12B | Q8_0 | 15.5 GB | ~56 tok/s |
| Gemma 3 12B | 12B | Q8_0 | 15.5 GB | ~56 tok/s |
| Step 3.5 Flash | 11B | Q8_0 | 14.3 GB | ~61 tok/s |
| Qwen3.5-9B | 9B | Q8_0 | 12.0 GB | ~75 tok/s |
| Llama 3.1 8B Instruct | 8B | Q8_0 | 10.9 GB | ~84 tok/s |
| Qwen3 VL 8B Instruct | 8B | Q8_0 | 10.9 GB | ~84 tok/s |
| Granite 4.1 8B | 8B | Q8_0 | 10.9 GB | ~84 tok/s |
Generating a token requires reading the model's active weights out of memory once. That makes decoding a bandwidth problem, not a compute problem — which is why a card's headline TFLOPS number tells you almost nothing about how fast it will feel.
At 896 GB/s, the RTX 5070 Ti can stream roughly 1,274 billion Q4 parameters per second. Divide that by a model's active parameter count and you have its ceiling in tokens per second. A 7B model lands near 182 tok/s; a 30B dense model near 42 tok/s. No amount of compute changes that ratio.
The buying rule this implies. Choose capacity for the model you want to run, then bandwidth for how it will feel. Two cards with identical VRAM can differ threefold in speed, and the spec sheet buries the number that decides it.
At 300W under sustained inference load, four hours a day works out to 1.2 kWh — about $6 a month at $0.17/kWh, or ₹288 at Indian residential rates. Size the PSU with headroom: 300W on the card means a 600W supply is the sensible floor once you add a CPU and transients.
15 of the models we track will not fit in 16 GB even at Q3_K_M. The first one you hit is GLM 4.7 Flash at 30B parameters, which needs 21 GB at Q4_K_M — 5 GB more than this card holds.
You are not completely stuck. --n-gpu-layers keeps as many layers on the GPU as fit and runs the remainder on the CPU, reading them across PCIe. It works, and it is dramatically slower: system RAM delivers perhaps 50–80 GB/s against this card's 896 GB/s, so every offloaded layer becomes a bottleneck. A model that is half-offloaded typically generates at walking pace. The usual conclusion is that a smaller model at a higher quantisation beats a larger one spilling into RAM.
| GPU | MSRP | VRAM | Bandwidth | Models it runs |
|---|---|---|---|---|
| RTX 5070 Ti | $749 | 16 GB | 896 GB/s | 38 |
| RTX 4070 Ti SUPER | $799 | 16 GB | 672 GB/s | 38 |
| RX 9070 XT | $599 | 16 GB | 645 GB/s | 38 |
| RTX 5070 | $549 | 12 GB | 672 GB/s | 26 |
Read that table by column, not by row. The VRAM column tells you which models are even on the menu; the bandwidth column tells you how they will feel. Cards that look interchangeable on price routinely differ by a factor of two in one column and not the other, and which of those matters depends entirely on whether your target model fits.
At Q4_K_M, 16 GB holds roughly 26B parameters. Mixture-of-experts models change the arithmetic for speed but not for capacity — all experts still have to be resident.
For models up to about 26B parameters at Q4, yes. Its 896 GB/s of bandwidth is what sets the speed ceiling, and its 16 GB of capacity is what sets the size ceiling. Beyond that you are offloading to system RAM and losing most of the benefit.
Its rated board power is 300W, and sustained generation will sit near that. Inference load is steadier than gaming load, so plan for the full figure continuously rather than in bursts.
For capacity, largely yes — llama.cpp and vLLM both split layers across cards, so two 16 GB cards hold roughly 32 GB of model. For speed, no: layers run sequentially across the split, so you gain very little throughput on a single request.
They answer different questions. VRAM decides whether a model runs at all; bandwidth decides how fast it generates once it does. Buy enough capacity for your target model first, because no amount of bandwidth rescues a model that does not fit.