Every page here answers one question with arithmetic rather than guesswork: will this model fit on your card, and how fast will it generate?
We track 53 open-weight models that fit consumer hardware. Each is sized across five quantisations, matched against every current GPU, and given a throughput estimate derived from memory bandwidth. Models are grouped below by the smallest card that holds them at Q4_K_M — find your VRAM, and everything in that group and above is within reach.
| Model | Parameters | VRAM at Q4 | Released |
|---|---|---|---|
| Qwen3.5-9B | 9B | 7.5 GB | 2026-03-10 |
| Llama 3.1 8B Instruct | 8B | 6.9 GB | 2024-07-23 |
| Qwen3 VL 8B Instruct | 8B | 6.9 GB | 2025-10-14 |
| Granite 4.1 8B | 8B | 6.9 GB | 2026-04-30 |
| Qwen3 8B | 8B | 6.9 GB | 2025-04-28 |
| Qwen3 VL 8B Thinking | 8B | 6.9 GB | 2025-10-14 |
| Granite 4.2 8B | 8B | 6.9 GB | 2026-08-31 |
| Qwen2.5 7B Instruct | 7B | 6.2 GB | 2024-10-16 |
| UI-TARS 7B | 7B | 6.2 GB | 2025-07-22 |
| Hy-MT2-7B | 7B | 6.2 GB | 2026-08-19 |
| Gemma 3 4B | 4B | 4.2 GB | 2025-03-13 |
| Nemotron 3.5 Content Safety | 4B | 4.2 GB | 2026-06-04 |
| Llama 3.2 3B Instruct | 3B | 3.5 GB | 2024-09-25 |
| Granite 4.0 Micro | 3B | 3.5 GB | 2025-10-20 |
| LFM2.5-2.6B | 3B | 3.3 GB | 2026-08-11 |
| Hy-MT2-1.8B | 2B | 2.7 GB | 2026-08-20 |
| Llama 3.2 1B Instruct | 1B | 2.1 GB | 2024-09-25 |
| Model | Parameters | VRAM at Q4 | Released |
|---|---|---|---|
| Qwen3 14B | 14B | 10.7 GB | 2025-04-28 |
| Ministral 3 14B 2512 | 14B | 10.7 GB | 2025-12-02 |
| Hunyuan A13B Instruct | 13B / 13B active | 10.1 GB | 2025-07-08 |
| Mistral Mistral Nemo | 12B | 9.5 GB | 2024-07-19 |
| Gemma 3 12B | 12B | 9.5 GB | 2025-03-13 |
| Step 3.5 Flash | 11B | 8.8 GB | 2026-01-29 |
| Model | Parameters | VRAM at Q4 | Released |
|---|---|---|---|
| gpt-oss-20b | 21B / 4B active | 13.7 GB | 2025-08-05 |
| Llama 4 Scout | 17B | 12.6 GB | 2025-04-05 |
| Llama 4 Maverick | 17B | 12.6 GB | 2025-04-05 |
| Model | Parameters | VRAM at Q4 | Released |
|---|---|---|---|
| Qwen3.6 35B A3B | 35B | 23.6 GB | 2026-04-27 |
| Qwen3.5-35B-A3B | 35B | 23.6 GB | 2026-02-25 |
| Laguna XS 2.1 | 33B | 22.4 GB | 2026-07-02 |
| Qwen2.5 Coder 32B Instruct | 32B | 21.8 GB | 2024-11-11 |
| Qwen3 VL 32B Instruct | 32B | 21.8 GB | 2025-10-23 |
| Qwen3 32B | 32B | 21.8 GB | 2025-04-28 |
| Gemma 4 31B | 31B | 21.2 GB | 2026-04-02 |
| GLM 4.7 Flash | 30B | 20.6 GB | 2026-01-19 |
| Qwen3 30B A3B | 30B | 20.6 GB | 2025-04-28 |
| Muse Glimmer 30B | 30B | 20.6 GB | 2026-08-09 |
| Qwen3 30B A3B Instruct 2507 | 30B / 3B active | 18.8 GB | 2025-07-29 |
| Qwen3 Coder 30B A3B Instruct | 30B | 20.6 GB | 2025-07-31 |
| Qwen3 VL 30B A3B Instruct | 30B | 20.6 GB | 2025-10-06 |
| Nemotron 3.5 Lightning | 30B / 3B active | 18.7 GB | 2026-08-11 |
| Nemotron 3 Nano Omni | 30B | 20.6 GB | 2026-04-28 |
| Nemotron 3 Nano 30B A3B | 30B | 20.6 GB | 2025-12-14 |
| Cohere North Mini Code | 30B / 3B active | 18.7 GB | 2026-06-17 |
| Qwen3.8 27B | 27B | 18.7 GB | 2026-08-14 |
| Qwen3.6 27B | 27B | 18.7 GB | 2026-04-27 |
| Qwen3.5-27B | 27B | 18.7 GB | 2026-02-25 |
| Gemma 2 27B | 27B | 18.7 GB | 2024-07-13 |
| Gemma 4 26B A4B | 25B | 17.6 GB | 2026-04-03 |
| Mistral Small 3.1 24B | 24B | 16.9 GB | 2025-03-17 |
| Voxtral Small 24B 2507 | 24B | 16.9 GB | 2025-10-30 |
| Mistral Small 3.2 24B | 24B | 16.9 GB | 2025-06-20 |
| Mistral Small 3 | 24B | 16.9 GB | 2025-01-30 |
| Model | Parameters | VRAM at Q4 | Released |
|---|---|---|---|
| Mixtral 8x22B Instruct | 39B / 39B active | 26.0 GB | 2024-04-17 |
Mixture-of-experts models appear in these groups by their total size, because that is what has to be resident in memory. Their active parameter count — the second figure — is what sets their speed, and it is often an order of magnitude smaller. A model listed at 284B that activates 13B per token will fill a card like a 284B model and run like a 13B one.