CohereLabs · text

c4ai-command-r-plus

CohereLabs/c4ai-command-r-plus

c4ai-command-r-plus at Q4_K_M is exactly 62,750,607,680 bytes (58.44 GiB / 62.75 GB) — an effective 4.836 bits per weight, not the nominal 4.

From the file· summed from 2 file(s)
Parameters
104B
Architecture
command-r
Context
native (config.json)
License
cc-by-nc-4.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S21.59 GiB23,181,871,5201.786dranger003
IQ1_S21.59 GiB23,181,871,5201.786qwp4w3hyb
IQ1_M23.49 GiB25,217,943,9681.943dranger003
IQ2_XXS26.65 GiB28,611,398,0482.205qwp4w3hyb
IQ2_XXS26.65 GiB28,611,398,0482.205dranger003
IQ2_XS29.46 GiB31,628,151,2002.437qwp4w3hyb
IQ2_XS29.46 GiB31,628,151,2002.437dranger003
IQ2_S31.04 GiB33,324,485,0242.568qwp4w3hyb
IQ2_S31.04 GiB33,324,485,0242.568dranger003
IQ2_M33.56 GiB36,039,248,2882.777dranger003
IQ2_M33.56 GiB36,039,248,2882.777qwp4w3hyb
Q2_K_S34.08 GiB36,595,452,3202.820dranger003
Q2_K36.78 GiB39,497,386,1123.044pmysl
Q2_K36.78 GiB39,497,386,4003.044dranger003
IQ3_XXS37.87 GiB40,658,749,8563.133dranger003
IQ3_XXS37.87 GiB40,658,749,8563.133qwp4w3hyb
IQ3_XS40.61 GiB43,599,415,7123.360dranger003
IQ3_XS40.61 GiB43,599,415,7123.360qwp4w3hyb
Q3_K_S42.70 GiB45,851,756,6723.534pmysl
IQ3_S42.80 GiB45,958,711,7123.542qwp4w3hyb
IQ3_S42.80 GiB45,958,711,7123.542dranger003
IQ3_M44.41 GiB47,683,357,0883.675qwp4w3hyb
IQ3_M44.41 GiB47,683,357,0883.675dranger003
Q3_K_M2 shards47.48 GiB50,982,439,2323.929pmysl
Q3_K_M2 shards47.48 GiB50,982,439,5523.929dranger003
Q3_K_L2 shards51.60 GiB55,402,187,0724.269pmysl
Q3_K_L2 shards51.60 GiB55,402,187,3924.269dranger003
IQ4_XS2 shards52.34 GiB56,201,202,3044.331dranger003
IQ4_NL2 shards55.25 GiB59,321,764,4804.572dranger003
Q4_K_S2 shards55.55 GiB59,642,628,4164.596pmysl
Q4_K_S2 shards55.55 GiB59,642,628,7044.596dranger003
Q4_K_M2 shards58.44 GiB62,750,607,6804.836pmysl
Q4_K_M2 shards58.44 GiB62,750,607,9684.836dranger003
Q5_K_S2 shards66.87 GiB71,804,012,8645.534pmysl
Q5_K_S2 shards66.87 GiB71,804,013,1845.534dranger003
Q5_K_M2 shards68.57 GiB73,622,243,6485.674pmysl
Q5_K_M2 shards68.57 GiB73,622,243,9685.674dranger003
Q6_K2 shards79.32 GiB85,173,356,8646.564pmysl
Q6_K2 shards79.32 GiB85,173,356,8646.564dranger003
Q8_03 shards102.74 GiB110,314,604,9928.501pmysl

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 54.38 GiB. The real file is 58.44 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does c4ai-command-r-plus need?
Q4_K_M is exactly 62,750,607,680 bytes (58.44 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of c4ai-command-r-plus should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.