google · text · gated

gemma-7b

google/gemma-7b

gemma-7b at Q4_K_M is exactly 5,329,758,592 bytes (4.96 GiB / 5.33 GB) — an effective 4.994 bits per weight, not the nominal 4. Its KV cache at 32K is 14.00 GiB.

From the file· summed from 1 file(s)From the file· KV from mirror (mirror:unsloth/gemma-7b)
Parameters
8.5B
Architecture
gemma
28 layers
Context
8,192
native (config.json)
License
gemma

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S2.01 GiB2,160,192,8002.024mradermacher
I1-IQ1_M2.16 GiB2,320,035,1042.174mradermacher
I1-IQ2_XXS2.41 GiB2,586,438,9442.424mradermacher
I1-IQ2_XS2.62 GiB2,810,572,0642.634mradermacher
I1-IQ2_S2.72 GiB2,918,903,0722.735mradermacher
I1-IQ2_M2.92 GiB3,132,026,1442.935mradermacher
Q2_K3.24 GiB3,481,446,7843.262MaziyarPanahi
I1-Q2_K3.24 GiB3,481,447,7123.262mradermacher
I1-IQ3_XXS3.25 GiB3,487,100,1923.268mradermacher
I1-IQ3_XS3.54 GiB3,800,739,1043.561mradermacher
Q3_K_S3.71 GiB3,982,403,9683.732MaziyarPanahi
I1-Q3_K_S3.71 GiB3,982,404,8963.732mradermacher
I1-IQ3_S3.71 GiB3,982,404,8963.732mradermacher
I1-IQ3_M3.82 GiB4,106,071,3283.848mradermacher
Q3_K_M4.07 GiB4,369,328,5124.094MaziyarPanahi
I1-Q3_K_M4.07 GiB4,369,329,4404.094mradermacher
Q3_K_L4.39 GiB4,709,067,1364.412MaziyarPanahi
I1-Q3_K_L4.39 GiB4,709,068,0644.412mradermacher
I1-IQ4_XS4.44 GiB4,769,623,3284.469mradermacher
I1-Q4_04.68 GiB5,026,000,1604.710mradermacher
Q4_K_S4.70 GiB5,046,446,4644.729MaziyarPanahi
I1-Q4_K_S4.70 GiB5,046,447,3924.729mradermacher
Q4_K_M4.96 GiB5,329,758,5924.994MaziyarPanahi
I1-Q4_K_M4.96 GiB5,329,759,5204.994mradermacher
Q5_K_S5.57 GiB5,980,727,6805.604MaziyarPanahi
I1-Q5_K_S5.57 GiB5,980,728,6085.604mradermacher
Q5_K_M5.72 GiB6,144,502,1445.758MaziyarPanahi
I1-Q5_K_M5.72 GiB6,144,503,0725.758mradermacher
Q6_K6.53 GiB7,010,167,1686.569MaziyarPanahi
I1-Q6_K6.53 GiB7,010,168,0966.569mradermacher
Q8_08.45 GiB9,077,844,3528.506MaziyarPanahi

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.75 GiB1.75 GiB28 / 0 / 0
8,1923.50 GiB3.50 GiB28 / 0 / 0
16,3847.00 GiB7.00 GiB28 / 0 / 0
32,76814.00 GiB14.00 GiB28 / 0 / 0
65,53628.00 GiB28.00 GiB28 / 0 / 0
131,07256.00 GiB56.00 GiB28 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.47 GiB. The real file is 4.96 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from mirror:unsloth/gemma-7b
Layers
28
Attention heads
16
KV heads
16
Head dim
256
Hidden size
3072
Vocab
256,000
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does gemma-7b need?
Q4_K_M is exactly 5,329,758,592 bytes (4.96 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is gemma-7b's KV cache?
14.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of gemma-7b should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.