Olmo-3.1-32B-Instruct
allenai/Olmo-3.1-32B-InstructOlmo-3.1-32B-Instruct at Q4_K_M is exactly 19,482,034,976 bytes (18.14 GiB / 19.48 GB) — an effective 4.835 bits per weight, not the nominal 4. Its KV cache at 32K is 2.84 GiB, not the 8.00 GiB a flat formula predicts.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| UD-IQ1_S | 6.75 GiB | 7,245,910,336 | 1.798 | — | unsloth |
| UD-IQ1_M | 7.33 GiB | 7,865,225,536 | 1.952 | — | unsloth |
| IQ2_XXS | 8.15 GiB | 8,756,307,456 | 2.173 | — | bartowski |
| UD-IQ2_XXS | 8.30 GiB | 8,914,620,736 | 2.212 | — | unsloth |
| IQ2_XS | 9.02 GiB | 9,685,607,936 | 2.404 | — | bartowski |
| IQ2_S | 9.40 GiB | 10,088,697,792 | 2.504 | — | bartowski |
| IQ2_M | 10.21 GiB | 10,965,569,472 | 2.721 | — | bartowski |
| UD-IQ2_M | 10.29 GiB | 11,047,571,776 | 2.742 | — | unsloth |
| Q2_K | 11.18 GiB | 12,005,941,632 | 2.980 | — | bartowski |
| Q2_K | 11.18 GiB | 12,005,942,016 | 2.980 | — | unsloth |
| Q2_K_L | 11.29 GiB | 12,126,275,616 | 3.010 | — | unsloth |
| Q2_K_L | 11.65 GiB | 12,507,331,616 | 3.104 | — | bartowski |
| IQ3_XXS | 11.68 GiB | 12,540,399,552 | 3.112 | — | bartowski |
| UD-IQ3_XXS | 11.78 GiB | 12,650,377,536 | 3.140 | — | unsloth |
| IQ3_XS | 12.45 GiB | 13,371,427,648 | 3.319 | — | bartowski |
| Q3_K_S | 13.09 GiB | 14,058,244,928 | 3.489 | — | bartowski |
| Q3_K_S | 13.09 GiB | 14,058,245,312 | 3.489 | — | unsloth |
| IQ3_M | 13.48 GiB | 14,476,036,928 | 3.593 | — | bartowski |
| Q3_K_M | 14.53 GiB | 15,600,962,368 | 3.872 | — | bartowski |
| Q3_K_M | 14.53 GiB | 15,600,962,752 | 3.872 | — | unsloth |
| Q3_K_L | 15.75 GiB | 16,912,993,088 | 4.198 | — | bartowski |
| IQ4_XS | 16.14 GiB | 17,332,139,232 | 4.302 | — | bartowski |
| IQ4_XS | 16.16 GiB | 17,348,184,096 | 4.306 | — | unsloth |
| IQ4_NL | 17.06 GiB | 18,312,873,632 | 4.545 | — | bartowski |
| IQ4_NL | 17.06 GiB | 18,312,874,016 | 4.545 | — | unsloth |
| Q4_0 | 17.08 GiB | 18,341,709,472 | 4.552 | — | bartowski |
| Q4_0 | 17.08 GiB | 18,341,709,856 | 4.552 | — | unsloth |
| Q4_K_S | 17.15 GiB | 18,415,109,792 | 4.570 | — | bartowski |
| Q4_K_S | 17.15 GiB | 18,415,110,176 | 4.570 | — | unsloth |
| Q4_K_M | 18.14 GiB | 19,482,034,976 | 4.835 | — | lmstudio-community |
| Q4_K_M | 18.14 GiB | 19,482,035,872 | 4.835 | — | bartowski |
| Q4_K_M | 18.14 GiB | 19,482,036,256 | 4.835 | — | unsloth |
| Q4_K_L | 18.50 GiB | 19,863,092,256 | 4.930 | — | bartowski |
| Q4_1 | 18.86 GiB | 20,253,370,912 | 5.027 | — | bartowski |
| Q4_1 | 18.86 GiB | 20,253,371,296 | 5.027 | — | unsloth |
| Q5_K_S | 20.71 GiB | 22,235,811,232 | 5.519 | — | bartowski |
| Q5_K_S | 20.71 GiB | 22,235,811,616 | 5.519 | — | unsloth |
| Q5_K_M | 21.29 GiB | 22,859,713,952 | 5.673 | — | bartowski |
| Q5_K_M | 21.29 GiB | 22,859,714,336 | 5.673 | — | unsloth |
| Q5_K_L | 21.58 GiB | 23,176,592,416 | 5.752 | — | bartowski |
KV cache by context
| Context | KV cache (f16)● | Flat formula | Overstated by | Full / windowed / recurrent |
|---|---|---|---|---|
| 4,096 | 1.00 GiB | 1.00 GiB | — | 16 / 48 / 0 |
| 8,192 | 1.34 GiB | 2.00 GiB | 1.49× | 16 / 48 / 0 |
| 16,384 | 1.84 GiB | 4.00 GiB | 2.17× | 16 / 48 / 0 |
| 32,768 | 2.84 GiB | 8.00 GiB | 2.81× | 16 / 48 / 0 |
| 65,536 | 4.84 GiB | 16.00 GiB | 3.30× | 16 / 48 / 0 |
| 131,072 | 8.84 GiB | 32.00 GiB | 3.62× | 16 / 48 / 0 |
48 of 64 layers cache only a 4,096-token window rather than the full context, on a period of . Figures assume the default configuration; --swa-full disables the saving entirely.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 16.89 GiB. The real file is 18.14 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 8.00 GiB at 32K context where the real figure is 2.84 GiB, because most of this model's layers cache a fixed window rather than the whole context.
Architecture
Questions people ask
- How much VRAM does Olmo-3.1-32B-Instruct need?
- Q4_K_M is exactly 19,482,034,976 bytes (18.14 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- How large is Olmo-3.1-32B-Instruct's KV cache?
- 2.84 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
- Which quantization of Olmo-3.1-32B-Instruct should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.