Quantization format

Q2_K

Files labelled Q2_K average 3.337 effective bits per weight across 2331 real quantizations — not the nominal 2.625. That is 27% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.

From the file· 2331 files measured
Nominal bpw
2.625
from the block layout
Measured average
3.337
2331 files
Range
1.078–19.476
varies by architecture
File sizes
0.01 GiB+
up to 864.81 GiB

Note

Blogs usually cite 2.5625. The block layout gives 2.625.

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
Qwen3.6-35B-A3B12.58 GiB3.006+14%
gemma-4-26B-A4B-it10.20 GiB3.301+26%
gemma-4-12B-it4.50 GiB3.231+23%
DeepSeek-V4-Flash92.86 GiB2.742+4%
Qwen3.5-4B2.04 GiB3.763+43%
Hy399.36 GiB2.857+9%
Qwen3-Coder-30B-A3B-Instruct10.49 GiB2.950+12%
gemma-4-31B-it11.76 GiB3.231+23%
Qwen3-VL-30B-A3B-Instruct10.49 GiB2.899+10%
FLUX.2-klein-9B3.71 GiB3.508+34%
LTX-2.37.39 GiB6.903+163%
Llama-3.2-1B-Instruct0.54 GiB3.760+43%
gemma-4-E2B-it2.81 GiB4.716+80%
gpt-oss-20b10.68 GiB4.265+62%
Qwen3-8B3.06 GiB3.205+22%
Qwen3.5-0.8B0.43 GiB4.252+62%
Qwen3-VL-8B-Instruct-abliterated-v13.06 GiB2.995+14%
UI-TARS-1.5-7B2.81 GiB2.910+11%
Qwen3-4B1.55 GiB3.320+26%
Llama-3.1-8B-Instruct2.96 GiB3.167+21%
Ornith-1.0-35B11.75 GiB2.912+11%
Llama-3.2-3B-Instruct1.27 GiB3.396+29%
ThinkingCap-Qwen3.6-27B11.03 GiB3.462+32%
whisper-medium0.25 GiB2.795+6%
Wan2.2-I2V-A14B4.94 GiB2.968+13%
From the filewhat these mean

Other quantizations