ibm-granite · text

granite-20b-code-instruct-8k

ibm-granite/granite-20b-code-instruct-8k

granite-20b-code-instruct-8k at Q4_K_M is exactly 12,820,208,544 bytes (11.94 GiB / 12.82 GB) — an effective 5.111 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
20.1B
Architecture
starcoder
52 layers
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S4.21 GiB4,517,697,5681.801bartowski
I1-IQ1_S4.21 GiB4,517,698,0161.801mradermacher
IQ1_S4.21 GiB4,517,698,4641.801legraphista
IQ1_M4.58 GiB4,912,437,2801.958bartowski
I1-IQ1_M4.58 GiB4,912,437,7281.958mradermacher
IQ1_M4.58 GiB4,912,438,1761.958legraphista
IQ2_XXS5.19 GiB5,570,336,8002.221bartowski
I1-IQ2_XXS5.19 GiB5,570,337,2482.221mradermacher
IQ2_XXS5.19 GiB5,570,337,6962.221legraphista
IQ2_XS5.74 GiB6,157,998,1122.455bartowski
I1-IQ2_XS5.74 GiB6,157,998,5602.455mradermacher
IQ2_XS5.74 GiB6,157,999,0082.455legraphista
IQ2_S6.08 GiB6,526,048,2882.602bartowski
I1-IQ2_S6.08 GiB6,526,048,7362.602mradermacher
IQ2_S6.08 GiB6,526,049,1842.602legraphista
I1-IQ2_M6.57 GiB7,052,368,3522.812mradermacher
IQ2_M6.57 GiB7,052,368,8002.812bartowski
IQ2_M6.57 GiB7,052,368,8002.812legraphista
Q2_K_S6.65 GiB7,145,020,3202.849legraphista
I1-Q2_K7.38 GiB7,929,485,7923.161mradermacher
Q2_K7.38 GiB7,929,486,2403.161legraphista
Q2_K7.38 GiB7,929,486,2403.161bartowski
Q2_K_L7.45 GiB8,002,624,4163.190bartowski
IQ3_XXS7.51 GiB8,062,539,8083.214bartowski
I1-IQ3_XXS7.51 GiB8,062,540,2563.214mradermacher
IQ3_XXS7.51 GiB8,062,540,7043.214legraphista
I1-IQ3_XS8.06 GiB8,658,557,4083.452mradermacher
IQ3_XS8.06 GiB8,658,557,8563.452legraphista
IQ3_XS8.06 GiB8,658,557,8563.452bartowski
IQ3_S8.32 GiB8,934,594,5923.562bartowski
I1-IQ3_S8.32 GiB8,934,595,0403.562mradermacher
I1-Q3_K_S8.32 GiB8,934,595,0403.562mradermacher
Q3_K_S8.32 GiB8,934,595,4883.562bartowski
IQ3_S8.32 GiB8,934,595,4883.562legraphista
Q3_K_S8.32 GiB8,934,595,4883.562legraphista
I1-IQ3_M8.93 GiB9,587,186,1443.822mradermacher
IQ3_M8.93 GiB9,587,186,5923.822bartowski
IQ3_M8.93 GiB9,587,186,5923.822legraphista
I1-Q3_K_M9.84 GiB10,566,293,9844.212mradermacher
Q3_K9.84 GiB10,566,294,4324.212legraphista

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 10.51 GiB. The real file is 11.94 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
52
Attention heads
KV heads
Head dim
Hidden size
Vocab
49,152
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does granite-20b-code-instruct-8k need?
Q4_K_M is exactly 12,820,208,544 bytes (11.94 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of granite-20b-code-instruct-8k should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.