Quantization format
Q4_0
Files labelled Q4_0 average 4.913 effective bits per weight across 1376 real quantizations — not the nominal 4.5. That is 9% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.
From the file· 1376 files measured
Nominal bpw
4.5
from the block layout
Measured average
4.913
1376 files
Range
1.303–32.453
varies by architecture
File sizes
0.02 GiB+
up to 549.26 GiB
Note
Legacy format, superseded by the K-quants for most uses.
What files with this label actually contain
tensor types across 34 parsed files
F32
10967
Q4_0
10113
F16
700
Q8_0
369
Q4_1
248
MXFP4
168
Q5_K
151
Q5_0
120
Q6_K
113
BF16
42
Q4_K
3
If Q4_0 were a uniform precision, this chart would have one bar. The F32 entries are normalization and bias tensors, which are never quantized; the higher K-quant entries are attention and output tensors deliberately promoted to protect quality.
Real files
one per model, most downloaded first
| Model | Size● | Effective bpw● | vs nominal● |
|---|---|---|---|
| Qwen3.6-27B | 14.71 GiB | 4.547 | +1% |
| embeddinggemma-300m | 0.26 GiB | 7.339 | +63% |
| Qwen3.6-35B-A3B | 20.51 GiB | 4.901 | +9% |
| Qwen3.5-9B | 5.01 GiB | 4.458 | +-1% |
| gemma-4-26B-A4B-it | 13.85 GiB | 4.482 | +-0% |
| gemma-4-12B-it | 6.28 GiB | 4.507 | +0% |
| Qwen3.5-4B | 2.41 GiB | 4.435 | +-1% |
| Hy3 | 158.62 GiB | 4.560 | +1% |
| gemma-4-E4B-it | 4.50 GiB | 4.838 | +8% |
| Qwen3-Coder-30B-A3B-Instruct | 16.19 GiB | 4.554 | +1% |
| gemma-4-31B-it | 16.15 GiB | 4.435 | +-1% |
| gemma-4-12B-it-qat-q4_0-unquantized | 6.50 GiB | 4.666 | +4% |
| Qwen3-VL-30B-A3B-Instruct | 16.19 GiB | 4.475 | +-1% |
| gemma-4-26B-A4B-it-qat-q4_0-unquantized | 13.45 GiB | 4.352 | +-3% |
| FLUX.2-klein-9B | 5.23 GiB | 4.949 | +10% |
| LTX-2.3 | 11.84 GiB | 11.061 | +146% |
| Llama-3.2-1B-Instruct | 0.72 GiB | 5.004 | +11% |
| gemma-4-E2B-it | 3.15 GiB | 5.276 | +17% |
| gpt-oss-20b | 10.71 GiB | 4.277 | +-5% |
| Qwen3-8B | 4.46 GiB | 4.676 | +4% |
| gemma-4-31B-it-qat-q4_0-unquantized | 16.44 GiB | 4.321 | +-4% |
| Qwen3.5-0.8B | 0.47 GiB | 4.645 | +3% |
| Qwen3-4B | 2.21 GiB | 4.725 | +5% |
| Llama-3.1-8B-Instruct | 4.35 GiB | 4.658 | +4% |
| Ornith-1.0-35B | 18.57 GiB | 4.603 | +2% |
●From the filewhat these mean