nbeerbower · text

Gemma4-Gutenberg-26B-A4B-lora

nbeerbower/Gemma4-Gutenberg-26B-A4B-lora

Gemma4-Gutenberg-26B-A4B-lora at Q4_K_M is exactly 16,796,016,000 bytes (15.64 GiB / 16.80 GB) — an effective 5.325 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
25.2B
Architecture
gemma4
Context
native (config.json)
License
gemma

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S7.72 GiB8,290,271,3922.628spiritfather
I1-IQ1_M8.07 GiB8,668,657,3122.748spiritfather
I1-IQ2_XXS8.66 GiB9,299,300,5122.948spiritfather
I1-IQ2_XS9.14 GiB9,816,430,7523.112spiritfather
I1-IQ2_S9.20 GiB9,873,201,3123.130spiritfather
I1-IQ2_M9.67 GiB10,377,715,8723.290spiritfather
Q2_K9.86 GiB10,582,737,2803.355spiritfather
I1-Q2_K9.86 GiB10,582,737,5683.355spiritfather
I1-Q2_K_S9.89 GiB10,624,481,9523.368spiritfather
I1-IQ3_XXS10.55 GiB11,325,694,1123.591spiritfather
I1-IQ3_XS10.84 GiB11,636,068,0003.689spiritfather
Q3_K_S11.38 GiB12,222,409,6003.875spiritfather
I1-Q3_K_S11.38 GiB12,222,409,8883.875spiritfather
I1-IQ3_S11.38 GiB12,222,409,8883.875spiritfather
I1-IQ3_M11.54 GiB12,392,563,8723.929spiritfather
Q3_K_M12.37 GiB13,286,733,6964.213spiritfather
I1-Q3_K_M12.37 GiB13,286,733,9844.213spiritfather
Q3_K_L12.88 GiB13,824,488,3204.383spiritfather
I1-Q3_K_L12.88 GiB13,824,488,6084.383spiritfather
I1-IQ4_XS12.96 GiB13,917,726,3684.412spiritfather
IQ4_XS13.10 GiB14,063,808,8964.459spiritfather
I1-Q4_013.49 GiB14,488,056,4804.593spiritfather
Q4_K_S14.40 GiB15,464,825,2164.903spiritfather
I1-Q4_K_S14.40 GiB15,464,825,5044.903spiritfather
I1-Q4_114.87 GiB15,969,576,6085.063spiritfather
Q4_K_M15.64 GiB16,796,016,0005.325spiritfather
I1-Q4_K_M15.64 GiB16,796,016,2885.325spiritfather
Q5_K_S16.75 GiB17,986,733,4405.703spiritfather
I1-Q5_K_S16.75 GiB17,986,733,7285.703spiritfather
Q5_K_M17.82 GiB19,132,890,4966.066spiritfather
I1-Q5_K_M17.82 GiB19,132,890,7846.066spiritfather
Q6_K21.08 GiB22,638,399,3607.177spiritfather
I1-Q6_K21.08 GiB22,638,399,6487.177spiritfather
Q8_025.02 GiB26,859,859,3288.516spiritfather

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 13.22 GiB. The real file is 15.64 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Gemma4-Gutenberg-26B-A4B-lora need?
Q4_K_M is exactly 16,796,016,000 bytes (15.64 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Gemma4-Gutenberg-26B-A4B-lora should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.