DavidAU · text

gemma-3-16b-it-BIG-G-GLM4.7-Flash-Valhalla-Heretic-Uncensored-Deep-Thinking

DavidAU/gemma-3-16b-it-BIG-G-GLM4.7-Flash-Valhalla-Heretic-Uncensored-Deep-Thinking

gemma-3-16b-it-BIG-G-GLM4.7-Flash-Valhalla-Heretic-Uncensored-Deep-Thinking at Q4_K_M is exactly 9,852,528,768 bytes (9.18 GiB / 9.85 GB) — an effective 4.919 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
16.0B
Architecture
gemma3
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S3.57 GiB3,834,571,3281.914mradermacher
I1-IQ1_M3.85 GiB4,138,422,8482.066mradermacher
I1-IQ2_XXS4.33 GiB4,644,842,0482.319mradermacher
I1-IQ2_XS4.73 GiB5,082,909,2482.538mradermacher
I1-IQ2_S4.97 GiB5,332,539,9682.662mradermacher
I1-IQ2_M5.34 GiB5,737,675,3282.864mradermacher
I1-Q2_K_S5.47 GiB5,874,811,3282.933mradermacher
Q2_K5.89 GiB6,326,118,5283.158mradermacher
I1-Q2_K5.89 GiB6,326,118,8483.158mradermacher
I1-IQ3_XXS5.96 GiB6,402,333,2483.196mradermacher
I1-IQ3_XS6.46 GiB6,938,798,5283.464mradermacher
Q3_K_S6.79 GiB7,289,374,8483.639mradermacher
I1-Q3_K_S6.79 GiB7,289,375,1683.639mradermacher
I1-IQ3_S6.79 GiB7,289,375,1683.639mradermacher
I1-IQ3_M7.04 GiB7,561,984,4483.775mradermacher
Q3_K_M7.50 GiB8,055,623,8084.022mradermacher
I1-Q3_K_M7.50 GiB8,055,624,1284.022mradermacher
Q3_K_L8.12 GiB8,715,735,1684.351mradermacher
I1-Q3_K_L8.12 GiB8,715,735,4884.351mradermacher
I1-IQ4_XS8.21 GiB8,814,531,0084.400mradermacher
IQ4_XS8.28 GiB8,888,258,6884.437mradermacher
I1-IQ4_NL8.65 GiB9,283,809,7284.635mradermacher
I1-Q4_08.67 GiB9,313,300,9284.649mradermacher
Q4_K_S8.70 GiB9,346,723,9684.666mradermacher
I1-Q4_K_S8.70 GiB9,346,724,2884.666mradermacher
Q4_K_M9.18 GiB9,852,528,7684.919mradermacher
I1-Q4_K_M9.18 GiB9,852,529,0884.919mradermacher
I1-Q4_19.52 GiB10,222,367,1685.103mradermacher
Q5_K_S10.39 GiB11,160,924,2885.572mradermacher
I1-Q5_K_S10.39 GiB11,160,924,6085.572mradermacher
Q5_K_M10.67 GiB11,453,900,9285.718mradermacher
I1-Q5_K_M10.67 GiB11,453,901,2485.718mradermacher
Q6_K12.25 GiB13,155,358,8486.567mradermacher
I1-Q6_K12.25 GiB13,155,359,1686.567mradermacher
Q8_015.87 GiB17,036,122,3688.505mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 8.39 GiB. The real file is 9.18 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does gemma-3-16b-it-BIG-G-GLM4.7-Flash-Valhalla-Heretic-Uncensored-Deep-Thinking need?
Q4_K_M is exactly 9,852,528,768 bytes (9.18 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of gemma-3-16b-it-BIG-G-GLM4.7-Flash-Valhalla-Heretic-Uncensored-Deep-Thinking should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.