CohereLabs · text

c4ai-command-r-plus-08-2024

CohereLabs/c4ai-command-r-plus-08-2024

c4ai-command-r-plus-08-2024 at Q4_K_M is exactly 62,750,608,128 bytes (58.44 GiB / 62.75 GB) — an effective 4.836 bits per weight, not the nominal 4. Its KV cache at 32K is 8.00 GiB.

From the file· summed from 2 file(s)From the file· KV from mirror (mirror:unsloth/c4ai-command-r-plus-08-2024)
Parameters
104B
Architecture
command-r
64 layers
Context
131,072
native (config.json)
License
cc-by-nc-4.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S21.59 GiB23,181,871,9361.786legraphista
IQ1_M23.49 GiB25,217,944,3841.943legraphista
IQ1_M23.49 GiB25,217,944,3841.943bartowski
IQ2_XXS26.65 GiB28,611,398,4642.205bartowski
IQ2_XXS26.65 GiB28,611,398,4642.205legraphista
IQ2_XS29.46 GiB31,628,151,6162.437bartowski
IQ2_XS29.46 GiB31,628,151,6162.437legraphista
IQ2_S31.04 GiB33,324,485,4402.568legraphista
IQ2_M33.56 GiB36,039,248,7042.777legraphista
IQ2_M33.56 GiB36,039,248,7042.777bartowski
Q2_K_S34.08 GiB36,595,452,7362.820legraphista
Q2_K36.78 GiB39,497,386,8163.044legraphista
Q2_K36.78 GiB39,497,386,8163.044bartowski
Q2_K_L37.49 GiB40,259,242,8163.103bartowski
IQ3_XXS37.87 GiB40,658,750,2723.133legraphista
IQ3_XXS37.87 GiB40,658,750,2723.133bartowski
IQ3_XS40.61 GiB43,599,416,1283.360legraphista
Q3_K_S42.70 GiB45,851,757,3763.534bartowski
Q3_K_S2 shards42.70 GiB45,851,757,5683.534legraphista
IQ3_S2 shards42.80 GiB45,958,712,3203.542legraphista
IQ3_M44.41 GiB47,683,357,5043.675bartowski
IQ3_M2 shards44.41 GiB47,683,357,6963.675legraphista
Q3_K_M2 shards47.48 GiB50,982,439,9363.929bartowski
Q3_K3 shards47.48 GiB50,982,440,0643.929legraphista
Q3_K_L2 shards51.60 GiB55,402,187,4884.269lmstudio-community
Q3_K_L2 shards51.60 GiB55,402,187,7764.269bartowski
Q3_K_L3 shards51.60 GiB55,402,187,9044.269legraphista
IQ4_XS2 shards52.34 GiB56,201,202,7204.331bartowski
IQ4_XS3 shards52.34 GiB56,201,202,8164.331legraphista
IQ4_NL3 shards55.25 GiB59,321,764,9924.572legraphista
Q4_02 shards55.35 GiB59,428,719,6484.580bartowski
Q4_K_S2 shards55.55 GiB59,642,629,1204.596bartowski
Q4_K_S3 shards55.55 GiB59,642,629,2484.596legraphista
Q4_K_M2 shards58.44 GiB62,750,608,1284.836lmstudio-community
Q4_K_M2 shards58.44 GiB62,750,608,4164.836bartowski
Q4_K3 shards58.44 GiB62,750,608,5124.836legraphista
Q4_K_L2 shards59.15 GiB63,512,464,3844.894bartowski
Q5_K_S3 shards66.87 GiB71,804,013,4085.534legraphista
Q5_K4 shards68.57 GiB73,622,244,3205.674legraphista
Q5_K_M2 shards68.57 GiB73,622,244,3845.674bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.00 GiB1.00 GiB64 / 0 / 0
8,1922.00 GiB2.00 GiB64 / 0 / 0
16,3844.00 GiB4.00 GiB64 / 0 / 0
32,7688.00 GiB8.00 GiB64 / 0 / 0
65,53616.00 GiB16.00 GiB64 / 0 / 0
131,07232.00 GiB32.00 GiB64 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 54.38 GiB. The real file is 58.44 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from mirror:unsloth/c4ai-command-r-plus-08-2024
Layers
64
Attention heads
96
KV heads
8
Head dim
128
Hidden size
12288
Vocab
256,000
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does c4ai-command-r-plus-08-2024 need?
Q4_K_M is exactly 62,750,608,128 bytes (58.44 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is c4ai-command-r-plus-08-2024's KV cache?
8.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of c4ai-command-r-plus-08-2024 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.