Runyard / Run locally

Run open-weight AI models locally

Every page here answers one question with arithmetic rather than guesswork: will this model fit on your card, and how fast will it generate?

We track 53 open-weight models that fit consumer hardware. Each is sized across five quantisations, matched against every current GPU, and given a throughput estimate derived from memory bandwidth. Models are grouped below by the smallest card that holds them at Q4_K_M — find your VRAM, and everything in that group and above is within reach.

Runs on 8GB

ModelParametersVRAM at Q4Released
Qwen3.5-9B9B7.5 GB2026-03-10
Llama 3.1 8B Instruct8B6.9 GB2024-07-23
Qwen3 VL 8B Instruct8B6.9 GB2025-10-14
Granite 4.1 8B8B6.9 GB2026-04-30
Qwen3 8B8B6.9 GB2025-04-28
Qwen3 VL 8B Thinking8B6.9 GB2025-10-14
Granite 4.2 8B8B6.9 GB2026-08-31
Qwen2.5 7B Instruct7B6.2 GB2024-10-16
UI-TARS 7B7B6.2 GB2025-07-22
Hy-MT2-7B7B6.2 GB2026-08-19
Gemma 3 4B4B4.2 GB2025-03-13
Nemotron 3.5 Content Safety4B4.2 GB2026-06-04
Llama 3.2 3B Instruct3B3.5 GB2024-09-25
Granite 4.0 Micro3B3.5 GB2025-10-20
LFM2.5-2.6B3B3.3 GB2026-08-11
Hy-MT2-1.8B2B2.7 GB2026-08-20
Llama 3.2 1B Instruct1B2.1 GB2024-09-25

Runs on 12GB

ModelParametersVRAM at Q4Released
Qwen3 14B14B10.7 GB2025-04-28
Ministral 3 14B 251214B10.7 GB2025-12-02
Hunyuan A13B Instruct13B / 13B active10.1 GB2025-07-08
Mistral Mistral Nemo12B9.5 GB2024-07-19
Gemma 3 12B12B9.5 GB2025-03-13
Step 3.5 Flash11B8.8 GB2026-01-29

Runs on 16GB

ModelParametersVRAM at Q4Released
gpt-oss-20b21B / 4B active13.7 GB2025-08-05
Llama 4 Scout17B12.6 GB2025-04-05
Llama 4 Maverick17B12.6 GB2025-04-05

Runs on 24GB

ModelParametersVRAM at Q4Released
Qwen3.6 35B A3B35B23.6 GB2026-04-27
Qwen3.5-35B-A3B35B23.6 GB2026-02-25
Laguna XS 2.133B22.4 GB2026-07-02
Qwen2.5 Coder 32B Instruct32B21.8 GB2024-11-11
Qwen3 VL 32B Instruct32B21.8 GB2025-10-23
Qwen3 32B32B21.8 GB2025-04-28
Gemma 4 31B31B21.2 GB2026-04-02
GLM 4.7 Flash30B20.6 GB2026-01-19
Qwen3 30B A3B30B20.6 GB2025-04-28
Muse Glimmer 30B30B20.6 GB2026-08-09
Qwen3 30B A3B Instruct 250730B / 3B active18.8 GB2025-07-29
Qwen3 Coder 30B A3B Instruct30B20.6 GB2025-07-31
Qwen3 VL 30B A3B Instruct30B20.6 GB2025-10-06
Nemotron 3.5 Lightning30B / 3B active18.7 GB2026-08-11
Nemotron 3 Nano Omni30B20.6 GB2026-04-28
Nemotron 3 Nano 30B A3B30B20.6 GB2025-12-14
Cohere North Mini Code30B / 3B active18.7 GB2026-06-17
Qwen3.8 27B27B18.7 GB2026-08-14
Qwen3.6 27B27B18.7 GB2026-04-27
Qwen3.5-27B27B18.7 GB2026-02-25
Gemma 2 27B27B18.7 GB2024-07-13
Gemma 4 26B A4B25B17.6 GB2026-04-03
Mistral Small 3.1 24B24B16.9 GB2025-03-17
Voxtral Small 24B 250724B16.9 GB2025-10-30
Mistral Small 3.2 24B24B16.9 GB2025-06-20
Mistral Small 324B16.9 GB2025-01-30

Runs on 32GB

ModelParametersVRAM at Q4Released
Mixtral 8x22B Instruct39B / 39B active26.0 GB2024-04-17

Mixture-of-experts models appear in these groups by their total size, because that is what has to be resident in memory. Their active parameter count — the second figure — is what sets their speed, and it is often an order of magnitude smaller. A model listed at 284B that activates 13B per token will fill a card like a 284B model and run like a 13B one.

Tools and other sections