Llama-4-Scout-17B-16E-Instruct-4bit
mlx-community/Llama-4-Scout-17B-16E-Instruct-4bitLlama-4-Scout-17B-16E-Instruct-4bit at Q4_K_M is exactly 22,491,854,144 bytes (20.95 GiB / 22.49 GB) — an effective 10.598 bits per weight, not the nominal 4.
Shipped quantizations
KV cache by context
This model declares a 8,192-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.
A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 8.89 GiB. The real file is 20.95 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Questions people ask
- How much VRAM does Llama-4-Scout-17B-16E-Instruct-4bit need?
- Q4_K_M is exactly 22,491,854,144 bytes (20.95 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Is Llama-4-Scout-17B-16E-Instruct-4bit a mixture-of-experts model?
- Yes — 16 experts, 1 routed per token. Every expert must be resident, but only the routed ones are read per token, which is why its memory requirement and its speed behave very differently.
- Which quantization of Llama-4-Scout-17B-16E-Instruct-4bit should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.