Quantization format

Q4_K_S

Files labelled Q4_K_S average 4.795 effective bits per weight across 2360 real quantizations — not the nominal 4.5. That is 7% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.

From the file· 2360 files measured
Nominal bpw
4.5
from the block layout
Measured average
4.795
2360 files
Range
1.303–27.527
varies by architecture
File sizes
0.02 GiB+
up to 557.32 GiB

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
Qwen3.6-27B14.77 GiB4.566+1%
Qwen3.6-35B-A3B20.01 GiB4.781+6%
Qwen3.5-9B5.02 GiB4.470+-1%
gemma-4-26B-A4B-it14.76 GiB4.775+6%
gemma-4-12B-it6.30 GiB4.525+1%
Qwen3.5-4B2.41 GiB4.447+-1%
Hy3163.41 GiB4.698+4%
gemma-4-E4B-it4.51 GiB4.847+8%
Qwen3-Coder-30B-A3B-Instruct16.26 GiB4.574+2%
gemma-4-31B-it16.20 GiB4.451+-1%
Qwen3-VL-30B-A3B-Instruct16.26 GiB4.495+-0%
FLUX.2-klein-9B5.43 GiB5.141+14%
LTX-2.312.07 GiB11.274+151%
Llama-3.2-1B-Instruct0.72 GiB5.021+12%
gemma-4-E2B-it2.83 GiB4.753+6%
gpt-oss-20b10.82 GiB4.321+-4%
Qwen3-8B4.47 GiB4.690+4%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP32.25 GiB9.970+122%
Qwen3.5-0.8B0.47 GiB4.654+3%
Qwen3-VL-8B-Instruct-abliterated-v14.47 GiB4.382+-3%
UI-TARS-1.5-7B4.15 GiB4.301+-4%
Qwen3-4B2.22 GiB4.740+5%
Llama-3.1-8B-Instruct4.37 GiB4.675+4%
Ornith-1.0-35B19.18 GiB4.753+6%
Llama-3.2-3B-Instruct1.80 GiB4.801+7%
From the filewhat these mean

Other quantizations