Runyard / GPUs / RTX 5090

RTX 5090 for local AI

The RTX 5090 pairs 32 GB of VRAM with 1,792 GB/s of memory bandwidth. Those two numbers, in that order, determine everything about what it can run and how quickly.

NVIDIA · Launched 2025-01 · 575W · $1,999 MSRP · Updated 2026-09-01

Of the 53 open-weight models we track, 53 fit entirely in this card's memory at some usable quantisation. The largest is Mixtral 8x22B Instruct at 39B parameters, which works because only 39B are active per token.

Models this card runs

Highest-fidelity quantisation that fits in 32 GB, with an 8K context window:

ModelParametersBest quantVRAM usedEst. speed
Mixtral 8x22B Instruct39B / 39B activeQ5_K_M31.9 GB~52 tok/s
Qwen3.6 35B A3B35BQ5_K_M28.8 GB~57 tok/s
Qwen3.5-35B-A3B35BQ5_K_M28.8 GB~57 tok/s
Laguna XS 2.133BQ6_K31.0 GB~53 tok/s
Qwen2.5 Coder 32B Instruct32BQ6_K30.2 GB~54 tok/s
Qwen3 VL 32B Instruct32BQ6_K30.2 GB~54 tok/s
Qwen3 32B32BQ6_K30.2 GB~54 tok/s
Gemma 4 31B31BQ6_K29.3 GB~56 tok/s
GLM 4.7 Flash30BQ6_K28.4 GB~58 tok/s
Qwen3 30B A3B30BQ6_K28.4 GB~58 tok/s
Muse Glimmer 30B30BQ6_K28.4 GB~58 tok/s
Qwen3 30B A3B Instruct 250730B / 3B activeQ6_K26.6 GB~527 tok/s
Qwen3 Coder 30B A3B Instruct30BQ6_K28.4 GB~58 tok/s
Qwen3 VL 30B A3B Instruct30BQ6_K28.4 GB~58 tok/s
Nemotron 3.5 Lightning30B / 3B activeQ6_K26.6 GB~579 tok/s
Nemotron 3 Nano Omni30BQ6_K28.4 GB~58 tok/s
Nemotron 3 Nano 30B A3B30BQ6_K28.4 GB~58 tok/s
Cohere North Mini Code30B / 3B activeQ6_K26.6 GB~579 tok/s
Qwen3.8 27B27BQ6_K25.8 GB~64 tok/s
Qwen3.6 27B27BQ6_K25.8 GB~64 tok/s
Qwen3.5-27B27BQ6_K25.8 GB~64 tok/s
Gemma 2 27B27BQ6_K25.8 GB~64 tok/s
Gemma 4 26B A4B25BQ8_030.2 GB~54 tok/s
Mistral Small 3.1 24B24BQ8_028.9 GB~56 tok/s
Voxtral Small 24B 250724BQ8_028.9 GB~56 tok/s

Why bandwidth matters more than you think

Generating a token requires reading the model's active weights out of memory once. That makes decoding a bandwidth problem, not a compute problem — which is why a card's headline TFLOPS number tells you almost nothing about how fast it will feel.

At 1,792 GB/s, the RTX 5090 can stream roughly 2,549 billion Q4 parameters per second. Divide that by a model's active parameter count and you have its ceiling in tokens per second. A 7B model lands near 364 tok/s; a 30B dense model near 85 tok/s. No amount of compute changes that ratio.

The buying rule this implies. Choose capacity for the model you want to run, then bandwidth for how it will feel. Two cards with identical VRAM can differ threefold in speed, and the spec sheet buries the number that decides it.

Power and running cost

At 575W under sustained inference load, four hours a day works out to 2.3 kWh — about $12 a month at $0.17/kWh, or ₹552 at Indian residential rates. Size the PSU with headroom: 575W on the card means a 850W supply is the sensible floor once you add a CPU and transients.

How it compares at this price

GPUMSRPVRAMBandwidthModels it runs
RTX 5090$1,99932 GB1,792 GB/s53
RTX 4090$1,59924 GB1,008 GB/s53
RTX 3090$1,49924 GB936 GB/s53
RTX 5080$99916 GB960 GB/s38

Read that table by column, not by row. The VRAM column tells you which models are even on the menu; the bandwidth column tells you how they will feel. Cards that look interchangeable on price routinely differ by a factor of two in one column and not the other, and which of those matters depends entirely on whether your target model fits.

Frequently asked questions

How many parameters can the RTX 5090 run?

At Q4_K_M, 32 GB holds roughly 54B parameters. Mixture-of-experts models change the arithmetic for speed but not for capacity — all experts still have to be resident.

Is the RTX 5090 good for local LLMs?

For models up to about 54B parameters at Q4, yes. Its 1,792 GB/s of bandwidth is what sets the speed ceiling, and its 32 GB of capacity is what sets the size ceiling. Beyond that you are offloading to system RAM and losing most of the benefit.

How much power does the RTX 5090 draw during inference?

Its rated board power is 575W, and sustained generation will sit near that. Inference load is steadier than gaming load, so plan for the full figure continuously rather than in bursts.

Do two GPUs double what I can run?

For capacity, largely yes — llama.cpp and vLLM both split layers across cards, so two 32 GB cards hold roughly 64 GB of model. For speed, no: layers run sequentially across the split, so you gain very little throughput on a single request.

Is more VRAM or more bandwidth better on the RTX 5090?

They answer different questions. VRAM decides whether a model runs at all; bandwidth decides how fast it generates once it does. Buy enough capacity for your target model first, because no amount of bandwidth rescues a model that does not fit.

Related