ibm-granite · text

granite-34b-code-base-8k

ibm-granite/granite-34b-code-base-8k

granite-34b-code-base-8k at Q4_K_M is exactly 21,383,628,160 bytes (19.92 GiB / 21.38 GB) — an effective 5.074 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
33.7B
Architecture
starcoder
88 layers
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S6.87 GiB7,377,962,6561.751bartowski
I1-IQ1_S6.87 GiB7,377,963,1361.751mradermacher
IQ1_M7.49 GiB8,042,989,2161.908bartowski
I1-IQ1_M7.49 GiB8,042,989,6961.908mradermacher
IQ2_XXS8.52 GiB9,151,366,8162.171bartowski
I1-IQ2_XXS8.52 GiB9,151,367,2962.171mradermacher
IQ2_XS9.45 GiB10,141,877,9202.406bartowski
I1-IQ2_XS9.45 GiB10,141,878,4002.406mradermacher
IQ2_S10.04 GiB10,777,708,1922.557bartowski
I1-IQ2_S10.04 GiB10,777,708,6722.557mradermacher
IQ2_M10.86 GiB11,664,410,2722.768bartowski
I1-IQ2_M10.86 GiB11,664,410,7522.768mradermacher
Q2_K12.21 GiB13,107,021,4723.110bartowski
I1-Q2_K12.21 GiB13,107,021,9523.110mradermacher
IQ3_XXS12.44 GiB13,359,957,6643.170bartowski
I1-IQ3_XXS12.44 GiB13,359,958,1443.170mradermacher
IQ3_XS13.36 GiB14,340,834,9763.403bartowski
I1-IQ3_XS13.36 GiB14,340,835,4563.403mradermacher
IQ3_S13.79 GiB14,807,975,5843.514bartowski
Q3_K_S13.79 GiB14,807,975,5843.514bartowski
I1-IQ3_S13.79 GiB14,807,976,0643.514mradermacher
I1-Q3_K_S13.79 GiB14,807,976,0643.514mradermacher
IQ3_M14.84 GiB15,929,329,3123.780bartowski
I1-IQ3_M14.84 GiB15,929,329,7923.780mradermacher
Q3_K_M16.36 GiB17,567,860,3844.168bartowski
I1-Q3_K_M16.36 GiB17,567,860,8644.168mradermacher
IQ4_XS16.95 GiB18,195,826,3364.317bartowski
I1-IQ4_XS16.95 GiB18,195,826,8164.317mradermacher
IQ4_NL17.92 GiB19,238,241,9524.565bartowski
I1-Q4_018.01 GiB19,342,051,4564.590mradermacher
Q4_K_S18.11 GiB19,445,860,0004.614bartowski
I1-Q4_K_S18.11 GiB19,445,860,4804.614mradermacher
Q3_K_L18.21 GiB19,549,669,0244.639bartowski
I1-Q3_K_L18.21 GiB19,549,669,5044.639mradermacher
Q4_K_M19.92 GiB21,383,628,1605.074ibm-granite
Q4_K_M19.92 GiB21,383,628,4485.074bartowski
I1-Q4_K_M19.92 GiB21,383,628,9285.074mradermacher
Q5_K_S21.80 GiB23,407,904,4165.554bartowski
I1-Q5_K_S21.80 GiB23,407,904,8965.554mradermacher
Q5_K_M23.05 GiB24,749,852,3205.873bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 17.66 GiB. The real file is 19.92 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
88
Attention heads
KV heads
Head dim
Hidden size
Vocab
49,152
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does granite-34b-code-base-8k need?
Q4_K_M is exactly 21,383,628,160 bytes (19.92 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of granite-34b-code-base-8k should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.