PatricHugi · text

gemma-4-26B-roleplay-v2-merged

PatricHugi/gemma-4-26B-roleplay-v2-merged

gemma-4-26B-roleplay-v2-merged at Q4_K_M is exactly 16,796,015,616 bytes (15.64 GiB / 16.80 GB) — an effective 5.325 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
25.2B
Architecture
gemma4
Context
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S7.72 GiB8,290,271,0082.628mradermacher
I1-IQ1_M8.07 GiB8,668,656,9282.748mradermacher
I1-IQ2_XXS8.66 GiB9,299,300,1282.948mradermacher
I1-IQ2_XS9.14 GiB9,816,430,3683.112mradermacher
I1-IQ2_S9.20 GiB9,873,200,9283.130mradermacher
I1-IQ2_M9.67 GiB10,377,715,4883.290mradermacher
Q2_K9.86 GiB10,582,736,8963.355mradermacher
I1-Q2_K9.86 GiB10,582,737,1843.355mradermacher
I1-Q2_K_S9.89 GiB10,624,481,5683.368mradermacher
I1-IQ3_XXS10.55 GiB11,325,693,7283.591mradermacher
I1-IQ3_XS10.84 GiB11,636,067,6163.689mradermacher
Q3_K_S11.38 GiB12,222,409,2163.875mradermacher
I1-IQ3_S11.38 GiB12,222,409,5043.875mradermacher
I1-Q3_K_S11.38 GiB12,222,409,5043.875mradermacher
I1-IQ3_M11.54 GiB12,392,563,4883.929mradermacher
Q3_K_M12.37 GiB13,286,733,3124.213mradermacher
I1-Q3_K_M12.37 GiB13,286,733,6004.213mradermacher
Q3_K_L12.88 GiB13,824,487,9364.383mradermacher
I1-Q3_K_L12.88 GiB13,824,488,2244.383mradermacher
I1-IQ4_XS12.96 GiB13,917,725,9844.412mradermacher
IQ4_XS13.10 GiB14,063,808,5124.459mradermacher
I1-Q4_013.49 GiB14,488,056,0964.593mradermacher
Q4_K_S14.40 GiB15,464,824,8324.903mradermacher
I1-Q4_K_S14.40 GiB15,464,825,1204.903mradermacher
I1-Q4_114.87 GiB15,969,576,2245.063mradermacher
Q4_K_M15.64 GiB16,796,015,6165.325mradermacher
I1-Q4_K_M15.64 GiB16,796,015,9045.325mradermacher
Q5_K_S16.75 GiB17,986,733,0565.703mradermacher
I1-Q5_K_S16.75 GiB17,986,733,3445.703mradermacher
Q5_K_M17.82 GiB19,132,890,1126.066mradermacher
I1-Q5_K_M17.82 GiB19,132,890,4006.066mradermacher
Q6_K21.08 GiB22,638,398,9767.177mradermacher
I1-Q6_K21.08 GiB22,638,399,2647.177mradermacher
Q8_025.02 GiB26,859,858,9448.516mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 13.22 GiB. The real file is 15.64 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does gemma-4-26B-roleplay-v2-merged need?
Q4_K_M is exactly 16,796,015,616 bytes (15.64 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of gemma-4-26B-roleplay-v2-merged should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.