Estimate any GGUF file’s disk and VRAM size before you download
Hugging Face repos often list a dozen GGUF variants of the same model — Q2_K, Q4_K_M, IQ3_XS — without a clear size on disk. This tool gives you a precise estimate before you start a multi-gigabyte download, plus VRAM at load and which GPUs the file fits.
GGUF file size
4.65GB
4764 MB on disk
Q4_K_M — Recommended default · best quality/size tradeoff
File size formula: (params × bpw ÷ 8) × 1.03 metadata factor. VRAM at load adds ~0.5 GB runtime overhead.
Fits on
8 GB
RTX 3070, RTX 4060
12 GB
RTX 4070, RTX 3060 12GB
16 GB
RTX 4060 Ti 16GB, RTX 4080
24 GB
RTX 3090, RTX 4090
32 GB
RTX 5090
48 GB
RTX 6000 Ada
80 GB
A100 80GB, H100
96 GB
MI300X, M3 Ultra
192 GB
MI300X (192GB), H200
How it works
1. Bits per weight (bpw) is the average bit budget each model parameter gets after quantization. Q4_K_M is 4.85 bpw, Q8_0 is 8.5 bpw, FP16 is 16.0 bpw. K-quants spend more bits on important weights and fewer on the rest.
2. File bytes = params × bpw ÷ 8, then × 1.03 to account for GGUF metadata overhead (tokenizer, tensor map, arch config).
3. VRAM at load = file size + ~0.5 GB for the CUDA / Metal runtime, allocator, and inference buffers.
4. KV cache is approximated as (params/1B) × context_k × 0.5 MB. Real KV cache size depends on attention heads and head dim — for frontier models (Llama 3.1, Qwen 2.5) this is conservative.
When you’d use this
Before a long Hugging Face download
Repos often show 10+ GGUF variants. Check size first so you don't pull 40 GB you can't even load.
Picking between Q4_K_M and Q5_K_M
See exactly how many GB the upgrade costs — useful when you're right at the edge of your VRAM budget.
Planning long-context use
Toggle the KV cache and step from 8K to 128K to see when context starts eating more memory than the model itself.
Related Runyard tools
VRAM Calculator →
Full VRAM picture including KV cache by context length and TurboQuant compression.
GGUF Variant Chooser →
Decide which quant suffix to download once you know the size budget.
Quantization Picker →
Reverse the question — what quant gives the best quality for the VRAM you have?
Deep-dive: How much VRAM do you really need?
FAQ
Approximately 42 GB. The math is 70B params × 4.85 bpw ÷ 8 = 42.4 GB, plus ~3% GGUF metadata. You need at least 48 GB of VRAM to load it comfortably, or two 24 GB GPUs with offload.
Both use 4-bit blocks, but Q4_K_M (medium) is 4.85 bpw and Q4_K_S (small) is 4.58 bpw. Q4_K_S saves about 5–6% file size and VRAM, at the cost of slightly more perplexity. Q4_K_M is the community-recommended default.
Roughly, but VRAM is always slightly higher than the file size. You add ~0.5 GB of runtime overhead (CUDA context, allocator, inference buffers), plus the KV cache which grows with context length. At 8K context the KV cache is small; at 128K it can equal or exceed the model itself.
For most users, Q4_K_M is the practical floor — it gives nearly the quality of Q5/Q6 at a much smaller size. Going below Q4 (Q3_K_M, Q2_K, IQ2_XS) is only worth it if you have no other choice, and only on big models (70B+) where the redundancy absorbs the quality hit.
The file size on disk is fixed — context length does not change it. But at runtime, the KV cache grows linearly with context. For an 8B model at 4K context KV is around 0.05 GB; at 128K it is around 1–2 GB. For 70B at 128K, KV can be 8+ GB.