Quantization format
F16
Files labelled F16 average 15.704 effective bits per weight across 1161 real quantizations — not the nominal 16. That is -2% more than the label implies, because a quantization is a mixture: some tensors are always kept at higher precision.
From the file· 1161 files measured
Nominal bpw
16
from the block layout
Measured average
15.704
1161 files
Range
1.028–32.396
varies by architecture
File sizes
0.04 GiB+
up to 1250.09 GiB
Note
Unquantized half precision.
What files with this label actually contain
tensor types across 39 parsed files
F32
15920
F16
13018
BF16
866
MXFP4
72
If F16 were a uniform precision, this chart would have one bar. The F32 entries are normalization and bias tensors, which are never quantized; the higher K-quant entries are attention and output tensors deliberately promoted to protect quality.
Real files
one per model, most downloaded first
| Model | Size● | Effective bpw● | vs nominal● |
|---|---|---|---|
| nemotron-3.5-asr-streaming-0.6b | 1.20 GiB | 16.130 | +1% |
| parakeet-unified-en-0.6b | 1.15 GiB | 16.032 | +0% |
| FLUX.2-klein-9B | 16.91 GiB | 16.000 | +0% |
| cohere-transcribe-03-2026 | 3.82 GiB | 15.903 | +-1% |
| LTX-2.3 | 39.15 GiB | — | — |
| Llama-3.2-1B-Instruct | 2.31 GiB | 16.052 | +0% |
| gpt-oss-20b | 12.85 GiB | 5.129 | +-68% |
| Qwopus3.6-35B-A3B-v1 | 66.19 GiB | 15.814 | +-1% |
| Qwen3-VL-8B-Instruct-abliterated-v1 | 15.26 GiB | 14.954 | +-7% |
| UI-TARS-1.5-7B | 14.19 GiB | 14.701 | +-8% |
| parakeet-tdt-0.6b-v3 | 1.17 GiB | 16.022 | +0% |
| Llama-3.2-3B-Instruct | 5.99 GiB | 16.020 | +0% |
| ThinkingCap-Qwen3.6-27B | 50.90 GiB | 15.984 | +-0% |
| whisper-medium | 1.44 GiB | 16.149 | +1% |
| Wan2.1-T2V-1.3B | 2.65 GiB | 16.016 | +0% |
| Bielik-11B-v3.0-Instruct | 20.80 GiB | 16.001 | +0% |
| Voxtral-Mini-4B-Realtime-2602 | 8.27 GiB | 16.036 | +0% |
| LFM2.5-1.2B-Instruct | 2.18 GiB | 16.018 | +0% |
| Qwen2.5-32B-Instruct | 61.04 GiB | 16.002 | +0% |
| gemma-3-1b-it | 1.87 GiB | 16.054 | +0% |
| Qwen2.5-Coder-7B-Instruct | 14.19 GiB | 16.007 | +0% |
| Qwen3-VL-4B-Instruct | 7.50 GiB | 14.514 | +-9% |
| Qwen2.5-7B-Instruct | 14.19 GiB | 16.007 | +0% |
| Qwen3-VL-2B-Instruct | 3.21 GiB | 12.963 | +-19% |
| Qwen2.5-1.5B-Instruct | 2.88 GiB | 16.032 | +0% |
●From the filewhat these mean