Quantization format

UD_IQ3_S

UD_IQ3_S has a nominal rate of bits per weight. We have no measured files carrying this label yet.

From the file· 32 files measured
Nominal bpw
from the block layout
Measured average
3.250
32 files
Range
2.903–4.560
varies by architecture
File sizes
3.33 GiB+
up to 390.04 GiB

Real files

one per model, most downloaded first
ModelSizeEffective bpwvs nominal
Qwen3.6-35B-A3B12.74 GiB3.043
gemma-4-26B-A4B-it10.51 GiB3.402
DeepSeek-V4-Flash109.25 GiB3.226
Qwen-AgentWorld-35B-A3B13.96 GiB3.458
GLM-5.2287.44 GiB3.278
Ornith-1.0-35B13.96 GiB3.458
Kimi-K2.7-Code390.04 GiB3.165
Qwen3.5-122B-A10B43.36 GiB2.978
Qwen3-Coder-Next27.65 GiB2.981
Qwen3.5-35B-A3B12.65 GiB3.023
Laguna-S-2.145.10 GiB3.296
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1617.53 GiB4.560
Step-3.7-Flash74.54 GiB3.180
LFM2.5-8B-A1B3.33 GiB3.374
Ornith-1.0-397B150.77 GiB3.264
Qwen3.5-397B-A17B136.32 GiB2.903
North-Mini-Code-1.011.89 GiB3.350
MiniMax-M2.777.87 GiB2.925
MiniMax-M3162.75 GiB3.274
NVIDIA-Nemotron-3-Super-120B-A12B-BF1652.74 GiB3.665
Mistral-Small-4-119B-260341.36 GiB2.975
Huihui-gemma-4-26B-A4B-it-abliterated10.45 GiB3.381
MiMo-V2.5106.98 GiB2.957
GLM-5.1260.39 GiB2.967
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16233.30 GiB3.575
From the filewhat these mean

Other quantizations