Does Wanabi-Gemma4-31B fit in 8GB of VRAM?
Not at these settings. No indexed quantization of Wanabi-Gemma4-31B fits 8GB card at any context we compute, with f16 KV. The smallest shipped quantization is 6.67 GiB in weights alone, against 7.44 GiB usable. CPU offload can still run it, slowly.
Every quantization at every context
| Quant | Weights● | 4K◐ | 8K◐ | 16K◐ | 32K◐ | 64K◐ | 128K◐ |
|---|---|---|---|---|---|---|---|
| BF16 | 57.20 GiB | 59.9 | 60.5 | 61.7 | 64.2 | 69.2 | 79.2 |
| Q8_0 | 30.39 GiB | 33.1 | 33.7 | 34.9 | 37.4 | 42.4 | 52.4 |
| I1-Q6_K | 23.47 GiB | 26.1 | 26.8 | 28.0 | 30.5 | 35.5 | 45.5 |
| Q6_K | 23.47 GiB | 26.1 | 26.8 | 28.0 | 30.5 | 35.5 | 45.5 |
| Q5_1 | 21.55 GiB | 24.2 | 24.9 | 26.1 | 28.6 | 33.6 | 43.6 |
| I1-Q5_K_M | 20.35 GiB | 23.0 | 23.6 | 24.9 | 27.4 | 32.4 | 42.4 |
| Q5_K_M | 20.35 GiB | 23.0 | 23.6 | 24.9 | 27.4 | 32.4 | 42.4 |
| I1-Q5_K_S | 19.85 GiB | 22.5 | 23.2 | 24.4 | 26.9 | 31.9 | 41.9 |
| Q5_0 | 19.85 GiB | 22.5 | 23.2 | 24.4 | 26.9 | 31.9 | 41.9 |
| Q5_K_S | 19.85 GiB | 22.5 | 23.2 | 24.4 | 26.9 | 31.9 | 41.9 |
| I1-Q4_1 | 18.14 GiB | 20.8 | 21.4 | 22.7 | 25.2 | 30.2 | 40.2 |
| Q4_1 | 18.14 GiB | 20.8 | 21.4 | 22.7 | 25.2 | 30.2 | 40.2 |
| I1-Q4_K_M | 17.40 GiB | 20.1 | 20.7 | 22.0 | 24.5 | 29.5 | 39.5 |
| Q4_K_M | 17.40 GiB | 20.1 | 20.7 | 22.0 | 24.5 | 29.5 | 39.5 |
| I1-Q4_K_S | 16.54 GiB | 19.2 | 19.8 | 21.1 | 23.6 | 28.6 | 38.6 |
| Q4_K_S | 16.54 GiB | 19.2 | 19.8 | 21.1 | 23.6 | 28.6 | 38.6 |
| I1-Q4_0 | 16.49 GiB | 19.2 | 19.8 | 21.0 | 23.5 | 28.5 | 38.5 |
| Q4_0 | 16.44 GiB | 19.1 | 19.7 | 21.0 | 23.5 | 28.5 | 38.5 |
| I1-IQ4_XS | 15.59 GiB | 18.3 | 18.9 | 20.1 | 22.6 | 27.6 | 37.6 |
| I1-Q3_K_L | 15.49 GiB | 18.2 | 18.8 | 20.0 | 22.5 | 27.5 | 37.5 |
| Q3_K_L | 15.49 GiB | 18.2 | 18.8 | 20.0 | 22.5 | 27.5 | 37.5 |
| I1-Q3_K_M | 14.24 GiB | 16.9 | 17.5 | 18.8 | 21.3 | 26.3 | 36.3 |
| Q3_K_M | 14.24 GiB | 16.9 | 17.5 | 18.8 | 21.3 | 26.3 | 36.3 |
| I1-IQ3_M | 13.43 GiB | 16.1 | 16.7 | 18.0 | 20.5 | 25.5 | 35.5 |
| I1-IQ3_S | 12.82 GiB | 15.5 | 16.1 | 17.4 | 19.9 | 24.9 | 34.9 |
| I1-Q3_K_S | 12.82 GiB | 15.5 | 16.1 | 17.4 | 19.9 | 24.9 | 34.9 |
| Q3_K_S | 12.82 GiB | 15.5 | 16.1 | 17.4 | 19.9 | 24.9 | 34.9 |
| I1-IQ3_XS | 12.17 GiB | 14.9 | 15.5 | 16.7 | 19.2 | 24.2 | 34.2 |
| I1-IQ3_XXS | 11.25 GiB | 13.9 | 14.6 | 15.8 | 18.3 | 23.3 | 33.3 |
| I1-Q2_K | 11.10 GiB | 13.8 | 14.4 | 15.7 | 18.2 | 23.2 | 33.2 |
| Q2_K | 11.10 GiB | 13.8 | 14.4 | 15.7 | 18.2 | 23.2 | 33.2 |
| I1-Q2_K_S | 10.22 GiB | 12.9 | 13.5 | 14.8 | 17.3 | 22.3 | 32.3 |
| I1-IQ2_M | 10.17 GiB | 12.8 | 13.5 | 14.7 | 17.2 | 22.2 | 32.2 |
| I1-IQ2_S | 9.46 GiB | 12.1 | 12.8 | 14.0 | 16.5 | 21.5 | 31.5 |
| I1-IQ2_XS | 8.88 GiB | 11.6 | 12.2 | 13.4 | 15.9 | 20.9 | 30.9 |
| I1-IQ2_XXS | 8.08 GiB | 10.8 | 11.4 | 12.6 | 15.1 | 20.1 | 30.1 |
| I1-IQ1_M | 7.20 GiB | 9.9 | 10.5 | 11.7 | 14.2 | 19.2 | 29.2 |
| I1-IQ1_S | 6.67 GiB | 9.3 | 10.0 | 11.2 | 13.7 | 18.7 | 28.7 |
Figures are GiB of total memory: weights plus KV cache plus compute buffer and backend overhead. Weights and KV are near-exact; the overhead term is modeled. Hover any cell for the breakdown.
Why there are no speeds on this page
A capacity is not a card. Whether a model fits depends only on memory, so every figure above holds for any 8GB accelerator. How fast it runs depends on memory bandwidth, which varies several-fold between cards of the same capacity — so putting a tokens-per-second number here would be inventing one. Pick a specific card from hardware and the speed column appears.
Why other calculators disagree
A parameters × bits ÷ 8 estimate ignores two things that dominate at long context. First, the weights themselves are not the nominal rate — quantizations are mixtures, so the real file is consistently larger than the label implies. Second, most of this model's layers cache only a 1,024-token window rather than the full context.