North-Mini-Code-1.0
CohereLabs/North-Mini-Code-1.0North-Mini-Code-1.0 at Q4_K_M is exactly 18,744,024,640 bytes (17.46 GiB / 18.74 GB) — an effective 4.919 bits per weight, not the nominal 4. Its KV cache at 32K is 1.13 GiB, not the 3.06 GiB a flat formula predicts.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| IQ2_XXS | 7.93 GiB | 8,512,945,728 | 2.234 | — | bartowski |
| UD-IQ1_M | 8.74 GiB | 9,379,799,136 | 2.462 | — | unsloth |
| IQ2_XS | 8.77 GiB | 9,418,915,392 | 2.472 | — | bartowski |
| IQ2_S | 8.95 GiB | 9,608,707,648 | 2.522 | — | bartowski |
| UD-IQ2_XXS | 9.11 GiB | 9,782,452,320 | 2.567 | — | unsloth |
| UD-IQ2_M | 9.19 GiB | 9,863,438,432 | 2.588 | — | unsloth |
| IQ2_M | 9.82 GiB | 10,549,280,320 | 2.768 | — | bartowski |
| Q2_K | 10.33 GiB | 11,090,304,576 | 2.910 | — | bartowski |
| Q2_K_L | 10.45 GiB | 11,220,328,000 | 2.945 | — | bartowski |
| UD-IQ3_XXS | 10.90 GiB | 11,708,375,136 | 3.073 | — | unsloth |
| UD-IQ3_S | 11.89 GiB | 12,765,339,744 | 3.350 | — | unsloth |
| IQ3_XXS | 12.10 GiB | 12,993,338,944 | 3.410 | — | bartowski |
| Q3_K_S | 12.63 GiB | 13,565,697,600 | 3.560 | — | bartowski |
| IQ3_XS | 13.23 GiB | 14,207,426,112 | 3.728 | — | bartowski |
| Q3_K_M | 13.23 GiB | 14,209,048,128 | 3.729 | 442 | bartowski |
| UD-Q3_K_M | 13.24 GiB | 14,213,013,600 | 3.730 | — | unsloth |
| Q3_K_L | 13.74 GiB | 14,758,436,416 | 3.873 | — | bartowski |
| IQ3_M | 13.84 GiB | 14,863,621,696 | 3.901 | — | bartowski |
| UD-IQ4_XS | 14.18 GiB | 15,230,132,320 | 3.997 | — | unsloth |
| UD-IQ4_NL | 14.47 GiB | 15,532,122,208 | 4.076 | — | unsloth |
| IQ4_XS | 15.44 GiB | 16,583,843,392 | 4.352 | 442 | bartowski |
| IQ4_NL | 16.30 GiB | 17,496,694,336 | 4.592 | — | bartowski |
| Q4_0 | 16.32 GiB | 17,521,204,800 | 4.598 | 442 | bartowski |
| UD-Q4_K_S | 16.81 GiB | 18,045,558,880 | 4.736 | — | unsloth |
| Q4_K_S | 16.81 GiB | 18,050,080,320 | 4.737 | — | bartowski |
| Q4_K_M | 17.46 GiB | 18,744,024,640 | 4.919 | 442 | bartowski |
| Q4_K_L | 17.58 GiB | 18,874,048,064 | 4.953 | — | bartowski |
| UD-Q4_K_M | 17.88 GiB | 19,203,186,784 | 5.040 | — | unsloth |
| Q4_1 | 17.97 GiB | 19,296,706,112 | 5.064 | — | bartowski |
| Q5_K_S | 19.70 GiB | 21,148,098,112 | 5.550 | — | bartowski |
| UD-Q5_K_S | 20.13 GiB | 21,619,105,888 | 5.673 | — | unsloth |
| Q5_K_M | 20.34 GiB | 21,845,253,696 | 5.733 | 442 | bartowski |
| Q5_K_L | 20.47 GiB | 21,975,277,120 | 5.767 | — | bartowski |
| UD-Q5_K_M | 21.37 GiB | 22,946,603,104 | 6.022 | — | unsloth |
| UD-Q6_K | 23.76 GiB | 25,513,517,152 | 6.696 | — | unsloth |
| Q6_K | 24.59 GiB | 26,402,856,512 | 6.929 | — | bartowski |
| Q6_K_L | 24.71 GiB | 26,532,879,936 | 6.963 | — | bartowski |
| Q8_0 | 30.21 GiB | 32,437,263,936 | 8.512 | 442 | bartowski |
| Q8_0 | 30.21 GiB | 32,437,264,480 | 8.512 | — | unsloth |
| BF162 shards | 56.81 GiB | 61,004,406,272 | 16.009 | — | bartowski |
KV cache by context
| Context | KV cache (f16)● | Flat formula | Overstated by | Full / windowed / recurrent |
|---|---|---|---|---|
| 4,096 | 0.38 GiB | 0.38 GiB | — | 13 / 36 / 0 |
| 8,192 | 0.52 GiB | 0.77 GiB | 1.47× | 13 / 36 / 0 |
| 16,384 | 0.72 GiB | 1.53 GiB | 2.12× | 13 / 36 / 0 |
| 32,768 | 1.13 GiB | 3.06 GiB | 2.71× | 13 / 36 / 0 |
| 65,536 | 1.94 GiB | 6.13 GiB | 3.15× | 13 / 36 / 0 |
| 131,072 | 3.57 GiB | 12.25 GiB | 3.43× | 13 / 36 / 0 |
36 of 49 layers cache only a 4,096-token window rather than the full context, on a period of 4. Figures assume the default configuration; --swa-full disables the saving entirely.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 15.97 GiB. The real file is 17.46 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 3.06 GiB at 32K context where the real figure is 1.13 GiB, because most of this model's layers cache a fixed window rather than the whole context.
Architecture
Questions people ask
- How much VRAM does North-Mini-Code-1.0 need?
- Q4_K_M is exactly 18,744,024,640 bytes (17.46 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- How large is North-Mini-Code-1.0's KV cache?
- 1.13 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
- Is North-Mini-Code-1.0 a mixture-of-experts model?
- Yes — 128 experts, 8 routed per token. Every expert must be resident, but only the routed ones are read per token, which is why its memory requirement and its speed behave very differently.
- Which quantization of North-Mini-Code-1.0 should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.