/GGUF File Size Estimator

GGUF File Size Estimator

Estimate any GGUF file’s disk and VRAM size before you download

P-18

Hugging Face repos often list a dozen GGUF variants of the same model — Q2_K, Q4_K_M, IQ3_XS — without a clear size on disk. This tool gives you a precise estimate before you start a multi-gigabyte download, plus VRAM at load and which GPUs the file fits.

GGUF file size

4.65GB

4764 MB on disk

Weights @ 4.85 bpw: 4.52 GBGGUF metadata (~3%): 0.14 GBVRAM at load: 5.15 GBTotal VRAM: 5.15 GB

Q4_K_MRecommended default · best quality/size tradeoff

File size formula: (params × bpw ÷ 8) × 1.03 metadata factor. VRAM at load adds ~0.5 GB runtime overhead.

Fits on

8 GB

RTX 3070, RTX 4060

12 GB

RTX 4070, RTX 3060 12GB

16 GB

RTX 4060 Ti 16GB, RTX 4080

24 GB

RTX 3090, RTX 4090

32 GB

RTX 5090

48 GB

RTX 6000 Ada

80 GB

A100 80GB, H100

96 GB

MI300X, M3 Ultra

192 GB

MI300X (192GB), H200

How it works

The math behind the file size.

1. Bits per weight (bpw) is the average bit budget each model parameter gets after quantization. Q4_K_M is 4.85 bpw, Q8_0 is 8.5 bpw, FP16 is 16.0 bpw. K-quants spend more bits on important weights and fewer on the rest.

2. File bytes = params × bpw ÷ 8, then × 1.03 to account for GGUF metadata overhead (tokenizer, tensor map, arch config).

3. VRAM at load = file size + ~0.5 GB for the CUDA / Metal runtime, allocator, and inference buffers.

4. KV cache is approximated as (params/1B) × context_k × 0.5 MB. Real KV cache size depends on attention heads and head dim — for frontier models (Llama 3.1, Qwen 2.5) this is conservative.

When you’d use this

Three honest use cases.

Before a long Hugging Face download

Repos often show 10+ GGUF variants. Check size first so you don't pull 40 GB you can't even load.

Picking between Q4_K_M and Q5_K_M

See exactly how many GB the upgrade costs — useful when you're right at the edge of your VRAM budget.

Planning long-context use

Toggle the KV cache and step from 8K to 128K to see when context starts eating more memory than the model itself.

Related Runyard tools

Other downloads worth pairing.

VRAM Calculator

Full VRAM picture including KV cache by context length and TurboQuant compression.

GGUF Variant Chooser

Decide which quant suffix to download once you know the size budget.

Quantization Picker

Reverse the question — what quant gives the best quality for the VRAM you have?

Deep-dive: How much VRAM do you really need?

FAQ

Frequently asked questions.

How big is a Q4_K_M GGUF file for Llama 3 70B?

Approximately 42 GB. The math is 70B params × 4.85 bpw ÷ 8 = 42.4 GB, plus ~3% GGUF metadata. You need at least 48 GB of VRAM to load it comfortably, or two 24 GB GPUs with offload.

What is the difference between Q4_K_M and Q4_K_S?

Both use 4-bit blocks, but Q4_K_M (medium) is 4.85 bpw and Q4_K_S (small) is 4.58 bpw. Q4_K_S saves about 5–6% file size and VRAM, at the cost of slightly more perplexity. Q4_K_M is the community-recommended default.

Does GGUF file size match VRAM usage?

Roughly, but VRAM is always slightly higher than the file size. You add ~0.5 GB of runtime overhead (CUDA context, allocator, inference buffers), plus the KV cache which grows with context length. At 8K context the KV cache is small; at 128K it can equal or exceed the model itself.

What is the smallest GGUF quant that still works well?

For most users, Q4_K_M is the practical floor — it gives nearly the quality of Q5/Q6 at a much smaller size. Going below Q4 (Q3_K_M, Q2_K, IQ2_XS) is only worth it if you have no other choice, and only on big models (70B+) where the redundancy absorbs the quality hit.

How does context length affect GGUF memory needs?

The file size on disk is fixed — context length does not change it. But at runtime, the KV cache grows linearly with context. For an 8B model at 4K context KV is around 0.05 GB; at 128K it is around 1–2 GB. For 70B at 128K, KV can be 8+ GB.