Quantization format

IQ4_XS

Files labelled IQ4_XS average 4.556 effective bits per weight across 1968 real quantizations — not the nominal 4.25. That is 7% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.

From the file· 1968 files measured
Nominal bpw
4.25
from the block layout
Measured average
4.556
1968 files
Range
1.206–27.605
varies by architecture
File sizes
0.03 GiB+
up to 510.00 GiB

Note

Importance-matrix quantization; needs a calibration pass to produce.

What files with this label actually contain

tensor types across 37 parsed files
F32
8840
IQ4_XS
8586
Q5_K
1094
Q8_0
538
Q6_K
285
IQ4_NL
188
MXFP4
96
Q4_K
49
Q5_1
26
Q4_0
16
BF16
6

If IQ4_XS were a uniform precision, this chart would have one bar. The F32 entries are normalization and bias tensors, which are never quantized; the higher K-quant entries are attention and output tensors deliberately promoted to protect quality.

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
Qwen3.6-27B14.38 GiB4.446+5%
embeddinggemma-300m0.28 GiB8.016+89%
Qwen3.6-35B-A3B18.35 GiB4.383+3%
Qwen3.5-9B4.81 GiB4.284+1%
gemma-4-26B-A4B-it13.23 GiB4.281+1%
gemma-4-12B-it5.94 GiB4.265+0%
Qwen3.5-4B2.31 GiB4.253+0%
Hy3148.22 GiB4.261+0%
gemma-4-E4B-it4.39 GiB4.718+11%
Qwen3-Coder-30B-A3B-Instruct15.25 GiB4.291+1%
gemma-4-31B-it15.25 GiB4.188+-1%
Qwen3-VL-30B-A3B-Instruct15.25 GiB4.217+-1%
Llama-3.2-1B-Instruct0.69 GiB4.811+13%
gemma-4-E2B-it2.78 GiB4.660+10%
Qwen3-8B4.27 GiB4.475+5%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP61.06 GiB18.880+344%
Qwen3.5-0.8B0.46 GiB4.512+6%
Qwen3-VL-8B-Instruct-abliterated-v14.28 GiB4.191+-1%
UI-TARS-1.5-7B3.96 GiB4.101+-4%
Qwen3-4B2.11 GiB4.516+6%
Llama-3.1-8B-Instruct4.14 GiB4.431+4%
Ornith-1.0-35B17.51 GiB4.341+2%
Qwopus3.6-27B-Coder14.47 GiB4.473+5%
Llama-3.2-3B-Instruct1.70 GiB4.555+7%
ThinkingCap-Qwen3.6-27B14.50 GiB4.553+7%
From the filewhat these mean

Other quantizations