Quantization format

Q4_K_M

Files labelled Q4_K_M average 5.130 effective bits per weight across 3293 real quantizations — not the nominal 4.5. That is 14% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.

From the file· 3293 files measured
Nominal bpw
4.5
from the block layout
Measured average
5.130
3293 files
Range
1.346–28.901
varies by architecture
File sizes
0.01 GiB+
up to 885.58 GiB

Note

The community default. Real files land near 4.9.

What files with this label actually contain

tensor types across 38 parsed files
F32
13855
Q4_K
10100
Q6_K
1690
Q8_0
1077
F16
940
Q5_K
376
BF16
123
Q5_0
93
MXFP4
72
Q4_0
8

If Q4_K_M were a uniform precision, this chart would have one bar. The F32 entries are normalization and bias tensors, which are never quantized; the higher K-quant entries are attention and output tensors deliberately promoted to protect quality.

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
Qwen3.6-27B15.41 GiB4.765+6%
Qwen3.6-35B-A3B19.71 GiB4.710+5%
Qwen3.5-9B5.24 GiB4.663+4%
gemma-4-26B-A4B-it15.87 GiB5.134+14%
gemma-4-12b-it6.87 GiB4.938+10%
nemotron-3.5-asr-streaming-0.6b0.46 GiB6.217+38%
parakeet-unified-en-0.6b0.44 GiB6.175+37%
Qwen3.5-4B2.52 GiB4.648+3%
Qwythos-9B-Claude-Mythos-5-1M10.73 GiB9.791+118%
Hy3169.81 GiB4.882+8%
gemma-4-e4b-it4.64 GiB4.980+11%
Qwen3-Coder-30B-A3B-Instruct17.28 GiB4.862+8%
gemma-4-31b-it17.40 GiB4.574+2%
HyperCLOVAX-SEED-Text-Instruct-1.5B1.06 GiB5.722+27%
Qwen3-VL-30B-A3B-Instruct17.28 GiB4.778+6%
FLUX.2-klein-9B5.50 GiB5.208+16%
cohere-transcribe-03-20261.45 GiB6.034+34%
LTX-2.316.54 GiB15.451+243%
Llama-3.2-1B-Instruct0.75 GiB5.229+16%
gemma-4-e2b-it3.22 GiB5.407+20%
gpt-oss-20b10.83 GiB4.323+-4%
Qwen3-8B4.68 GiB4.911+9%
Qwopus3.6-35B-A3B-v120.22 GiB4.832+7%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP34.04 GiB10.524+134%
Qwen3.5-0.8B0.49 GiB4.832+7%
From the filewhat these mean

Other quantizations