ibm-granite · text

granite-20b-code-base-8k

ibm-granite/granite-20b-code-base-8k

granite-20b-code-base-8k at Q4_K_M is exactly 12,820,207,360 bytes (11.94 GiB / 12.82 GB) — an effective 5.111 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
20.1B
Architecture
starcoder
52 layers
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S4.21 GiB4,517,698,0481.801mradermacher
I1-IQ1_M4.58 GiB4,912,437,7601.958mradermacher
I1-IQ2_XXS5.19 GiB5,570,337,2802.221mradermacher
I1-IQ2_XS5.74 GiB6,157,998,5922.455mradermacher
I1-IQ2_S6.08 GiB6,526,048,7682.602mradermacher
I1-IQ2_M6.57 GiB7,052,368,3842.812mradermacher
Q2_K7.38 GiB7,929,485,5683.161mradermacher
I1-Q2_K7.38 GiB7,929,485,8243.161mradermacher
I1-IQ3_XXS7.51 GiB8,062,540,2883.214mradermacher
IQ3_XS8.06 GiB8,658,557,1843.452mradermacher
I1-IQ3_XS8.06 GiB8,658,557,4403.452mradermacher
Q3_K_S8.32 GiB8,934,594,8163.562mradermacher
IQ3_S8.32 GiB8,934,594,8163.562mradermacher
I1-Q3_K_S8.32 GiB8,934,595,0723.562mradermacher
I1-IQ3_S8.32 GiB8,934,595,0723.562mradermacher
IQ3_M8.93 GiB9,587,185,9203.822mradermacher
I1-IQ3_M8.93 GiB9,587,186,1763.822mradermacher
Q3_K_M9.84 GiB10,566,293,7604.212mradermacher
I1-Q3_K_M9.84 GiB10,566,294,0164.212mradermacher
I1-IQ4_XS10.19 GiB10,936,506,8804.360mradermacher
IQ4_XS10.32 GiB11,078,064,3844.416mradermacher
I1-Q4_010.81 GiB11,609,102,8484.628mradermacher
Q4_K_S10.86 GiB11,665,725,6964.651mradermacher
I1-Q4_K_S10.86 GiB11,665,725,9524.651mradermacher
Q3_K_L10.93 GiB11,736,504,5764.679mradermacher
I1-Q3_K_L10.93 GiB11,736,504,8324.679mradermacher
Q4_K_M11.94 GiB12,820,207,3605.111ibm-granite
Q4_K_M11.94 GiB12,820,207,8725.111mradermacher
I1-Q4_K_M11.94 GiB12,820,208,1285.111mradermacher
Q5_K_S13.05 GiB14,016,370,9445.588mradermacher
I1-Q5_K_S13.05 GiB14,016,371,2005.588mradermacher
Q5_K_M13.79 GiB14,809,340,1605.904mradermacher
I1-Q5_K_M13.79 GiB14,809,340,4165.904mradermacher
Q6_K15.49 GiB16,634,255,6166.631mradermacher
I1-Q6_K15.49 GiB16,634,255,8726.631mradermacher
Q8_020.01 GiB21,481,183,4888.564mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 10.51 GiB. The real file is 11.94 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
52
Attention heads
KV heads
Head dim
Hidden size
Vocab
49,152
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does granite-20b-code-base-8k need?
Q4_K_M is exactly 12,820,207,360 bytes (11.94 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of granite-20b-code-base-8k should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.