Runyard / GPUs / RTX 5070

RTX 5070 for local AI

The RTX 5070 pairs 12 GB of VRAM with 672 GB/s of memory bandwidth. Those two numbers, in that order, determine everything about what it can run and how quickly.

NVIDIA · Launched 2025-03 · 250W · $549 MSRP · Updated 2026-09-01

Of the 53 open-weight models we track, 26 fit entirely in this card's memory at some usable quantisation. The largest is gpt-oss-20b at 21B parameters, which works because only 4B are active per token.

Models this card runs

Highest-fidelity quantisation that fits in 12 GB, with an 8K context window:

ModelParametersBest quantVRAM usedEst. speed
gpt-oss-20b21B / 4B activeQ3_K_M11.1 GB~341 tok/s
Llama 4 Scout17BQ3_K_M10.5 GB~72 tok/s
Llama 4 Maverick17BQ3_K_M10.5 GB~72 tok/s
Qwen3 14B14BQ4_K_M10.7 GB~68 tok/s
Ministral 3 14B 251214BQ4_K_M10.7 GB~68 tok/s
Hunyuan A13B Instruct13B / 13B activeQ4_K_M10.1 GB~74 tok/s
Mistral Mistral Nemo12BQ5_K_M11.3 GB~63 tok/s
Gemma 3 12B12BQ5_K_M11.3 GB~63 tok/s
Step 3.5 Flash11BQ6_K11.7 GB~59 tok/s
Qwen3.5-9B9BQ6_K9.9 GB~72 tok/s
Llama 3.1 8B Instruct8BQ8_010.9 GB~63 tok/s
Qwen3 VL 8B Instruct8BQ8_010.9 GB~63 tok/s
Granite 4.1 8B8BQ8_010.9 GB~63 tok/s
Qwen3 8B8BQ8_010.9 GB~63 tok/s
Qwen3 VL 8B Thinking8BQ8_010.9 GB~63 tok/s
Granite 4.2 8B8BQ8_010.9 GB~63 tok/s
Qwen2.5 7B Instruct7BQ8_09.7 GB~72 tok/s
UI-TARS 7B7BQ8_09.7 GB~72 tok/s
Hy-MT2-7B7BQ8_09.7 GB~72 tok/s
Gemma 3 4B4BQ8_06.2 GB~126 tok/s
Nemotron 3.5 Content Safety4BQ8_06.2 GB~126 tok/s
Llama 3.2 3B Instruct3BQ8_05.0 GB~169 tok/s
Granite 4.0 Micro3BQ8_05.0 GB~169 tok/s
LFM2.5-2.6B3BQ8_04.6 GB~195 tok/s
Hy-MT2-1.8B2BQ8_03.6 GB~281 tok/s

Why bandwidth matters more than you think

Generating a token requires reading the model's active weights out of memory once. That makes decoding a bandwidth problem, not a compute problem — which is why a card's headline TFLOPS number tells you almost nothing about how fast it will feel.

At 672 GB/s, the RTX 5070 can stream roughly 956 billion Q4 parameters per second. Divide that by a model's active parameter count and you have its ceiling in tokens per second. A 7B model lands near 137 tok/s; a 30B dense model near 32 tok/s. No amount of compute changes that ratio.

The buying rule this implies. Choose capacity for the model you want to run, then bandwidth for how it will feel. Two cards with identical VRAM can differ threefold in speed, and the spec sheet buries the number that decides it.

Power and running cost

At 250W under sustained inference load, four hours a day works out to 1.0 kWh — about $5 a month at $0.17/kWh, or ₹240 at Indian residential rates. Size the PSU with headroom: 250W on the card means a 550W supply is the sensible floor once you add a CPU and transients.

Where this card runs out

27 of the models we track will not fit in 12 GB even at Q3_K_M. The first one you hit is Mistral Small 3.1 24B at 24B parameters, which needs 17 GB at Q4_K_M — 5 GB more than this card holds.

You are not completely stuck. --n-gpu-layers keeps as many layers on the GPU as fit and runs the remainder on the CPU, reading them across PCIe. It works, and it is dramatically slower: system RAM delivers perhaps 50–80 GB/s against this card's 672 GB/s, so every offloaded layer becomes a bottleneck. A model that is half-offloaded typically generates at walking pace. The usual conclusion is that a smaller model at a higher quantisation beats a larger one spilling into RAM.

How it compares at this price

GPUMSRPVRAMBandwidthModels it runs
RTX 5070$54912 GB672 GB/s26
RX 9070 XT$59916 GB645 GB/s38
RTX 4060 Ti 16GB$44916 GB288 GB/s38
RTX 5060 Ti 16GB$42916 GB448 GB/s38

Read that table by column, not by row. The VRAM column tells you which models are even on the menu; the bandwidth column tells you how they will feel. Cards that look interchangeable on price routinely differ by a factor of two in one column and not the other, and which of those matters depends entirely on whether your target model fits.

Frequently asked questions

How many parameters can the RTX 5070 run?

At Q4_K_M, 12 GB holds roughly 19B parameters. Mixture-of-experts models change the arithmetic for speed but not for capacity — all experts still have to be resident.

Is the RTX 5070 good for local LLMs?

For models up to about 19B parameters at Q4, yes. Its 672 GB/s of bandwidth is what sets the speed ceiling, and its 12 GB of capacity is what sets the size ceiling. Beyond that you are offloading to system RAM and losing most of the benefit.

How much power does the RTX 5070 draw during inference?

Its rated board power is 250W, and sustained generation will sit near that. Inference load is steadier than gaming load, so plan for the full figure continuously rather than in bursts.

Do two GPUs double what I can run?

For capacity, largely yes — llama.cpp and vLLM both split layers across cards, so two 12 GB cards hold roughly 24 GB of model. For speed, no: layers run sequentially across the split, so you gain very little throughput on a single request.

Is more VRAM or more bandwidth better on the RTX 5070?

They answer different questions. VRAM decides whether a model runs at all; bandwidth decides how fast it generates once it does. Buy enough capacity for your target model first, because no amount of bandwidth rescues a model that does not fit.

Related