Quantization format
Q5_K_S
Files labelled Q5_K_S average 5.656 effective bits per weight across 2319 real quantizations — not the nominal 5.5. That is 3% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.
From the file· 2319 files measured
Nominal bpw
5.5
from the block layout
Measured average
5.656
2319 files
Range
1.558–24.887
varies by architecture
File sizes
0.02 GiB+
up to 658.45 GiB
Real files
one per model, most downloaded first
| Model | Size● | Effective bpw● | vs nominal● |
|---|---|---|---|
| Qwen3.6-27B | 17.66 GiB | 5.459 | +-1% |
| Qwen3.6-35B-A3B | 23.33 GiB | 5.574 | +1% |
| Qwen3.5-9B | 5.92 GiB | 5.272 | +-4% |
| gemma-4-26B-A4B-it | 16.88 GiB | 5.464 | +-1% |
| gemma-4-12B-it | 7.64 GiB | 5.488 | +-0% |
| Qwen3.5-4B | 2.82 GiB | 5.193 | +-6% |
| Hy3 | 191.88 GiB | 5.517 | +0% |
| gemma-4-E4B-it | 5.03 GiB | 5.407 | +-2% |
| Qwen3-Coder-30B-A3B-Instruct | 19.63 GiB | 5.524 | +0% |
| gemma-4-31B-it | 19.67 GiB | 5.404 | +-2% |
| Qwen3-VL-30B-A3B-Instruct | 19.63 GiB | 5.428 | +-1% |
| FLUX.2-klein-9B | 6.46 GiB | 6.114 | +11% |
| LTX-2.3 | 14.01 GiB | 13.085 | +138% |
| Llama-3.2-1B-Instruct | 0.83 GiB | 5.778 | +5% |
| gemma-4-E2B-it | 3.09 GiB | 5.186 | +-6% |
| gpt-oss-20b | 10.91 GiB | 4.356 | +-21% |
| Qwen3-8B | 5.33 GiB | 5.588 | +2% |
| Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP | 38.01 GiB | 11.752 | +114% |
| Qwen3.5-0.8B | 0.53 GiB | 5.211 | +-5% |
| Qwen3-VL-8B-Instruct-abliterated-v1 | 5.33 GiB | 5.220 | +-5% |
| UI-TARS-1.5-7B | 4.95 GiB | 5.128 | +-7% |
| Qwen3-4B | 2.63 GiB | 5.616 | +2% |
| Llama-3.1-8B-Instruct | 5.21 GiB | 5.578 | +1% |
| Ornith-1.0-35B | 22.50 GiB | 5.575 | +1% |
| Llama-3.2-3B-Instruct | 2.11 GiB | 5.651 | +3% |
●From the filewhat these mean