Quantization format
Q4_K
Q4_K has a nominal rate of — bits per weight. We have no measured files carrying this label yet.
From the file· 164 files measured
Nominal bpw
—
from the block layout
Measured average
6.044
164 files
Range
3.201–28.888
varies by architecture
File sizes
0.02 GiB+
up to 165.75 GiB
Real files
one per model, most downloaded first
| Model | Size● | Effective bpw● | vs nominal● |
|---|---|---|---|
| embeddinggemma-300m | 0.28 GiB | 8.082 | — |
| nemotron-3.5-asr-streaming-0.6b | 0.38 GiB | 5.123 | — |
| DeepSeek-V4-Flash | 153.33 GiB | 4.527 | — |
| Qwythos-9B-Claude-Mythos-5-1M | 5.38 GiB | 4.914 | — |
| FLUX.2-klein-9B | 5.32 GiB | 5.038 | — |
| cohere-transcribe-03-2026 | 1.41 GiB | 5.849 | — |
| parakeet-tdt-0.6b-v3 | 0.39 GiB | 5.332 | — |
| whisper-medium | 0.41 GiB | 4.655 | — |
| Voxtral-Mini-4B-Realtime-2602 | 2.35 GiB | 4.559 | — |
| whisper-large-v3 | 0.83 GiB | 4.609 | — |
| whisper-large-v3-turbo | 0.44 GiB | 4.688 | — |
| gemma-4-E2B-it-qat-q4_0-unquantized | 3.18 GiB | 5.354 | — |
| GLM-4.7-Flash | 16.99 GiB | 4.675 | — |
| Qwen3-ASR-1.7B | 1.39 GiB | 5.077 | — |
| Z-Image-Turbo | 3.60 GiB | 5.023 | — |
| Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled | 20.22 GiB | 4.831 | — |
| Qwen3-ASR-0.6B | 0.59 GiB | 5.382 | — |
| Wan2.2-Animate-14B | 10.80 GiB | 5.372 | — |
| parakeet-tdt-0.6b-v2 | 0.37 GiB | 5.139 | — |
| Mistral-7B-Instruct-v0.3 | 4.07 GiB | 4.827 | — |
| Qwen3-TTS-12Hz-0.6B-Base | 0.50 GiB | 4.661 | — |
| gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinking | 17.40 GiB | 4.780 | — |
| canary-1b-v2 | 0.37 GiB | 3.201 | — |
| FLUX.1-Fill-dev | 6.45 GiB | 4.657 | — |
| Llama-3.1-8B | 4.58 GiB | 4.902 | — |
●From the filewhat these mean