google · text

translategemma-12b-it

google/translategemma-12b-it

translategemma-12b-it at Q4_K_M is exactly 7,300,776,544 bytes (6.80 GiB / 7.30 GB) — an effective 4.427 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
13.2B
Architecture
gemma3
Context
native (config.json)
License
gemma

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S2.75 GiB2,947,430,5281.787mradermacher
I1-IQ1_M2.95 GiB3,164,743,8081.919mradermacher
I1-IQ2_XXS3.28 GiB3,526,932,6082.139mradermacher
I1-IQ2_XS3.58 GiB3,840,276,6082.329mradermacher
I1-IQ2_S3.74 GiB4,020,725,8882.438mradermacher
I1-IQ2_M4.01 GiB4,310,476,9282.614mradermacher
I1-Q2_K_S4.14 GiB4,448,626,6882.697mradermacher
Q2_K4.44 GiB4,768,237,3122.891mradermacher
I1-Q2_K4.44 GiB4,768,237,5682.891mradermacher
I1-IQ3_XXS4.46 GiB4,784,916,6082.901mradermacher
I1-IQ3_XS4.85 GiB5,206,181,8883.157mradermacher
Q3_K_S5.08 GiB5,458,331,3923.309mradermacher
I1-Q3_K_S5.08 GiB5,458,331,6483.309mradermacher
I1-IQ3_S5.08 GiB5,458,331,6483.309mradermacher
I1-IQ3_M5.27 GiB5,655,738,3683.429mradermacher
Q3_K_M5.60 GiB6,008,833,7923.643mradermacher
I1-Q3_K_M5.60 GiB6,008,834,0483.643mradermacher
Q3_K_L6.04 GiB6,480,201,4723.929mradermacher
I1-Q3_K_L6.04 GiB6,480,201,7283.929mradermacher
I1-IQ4_XS6.10 GiB6,550,980,6083.972mradermacher
IQ4_XS6.15 GiB6,606,276,3524.006mradermacher
I1-IQ4_NL6.41 GiB6,887,180,2884.176mradermacher
I1-Q4_06.43 GiB6,909,298,6884.189mradermacher
Q4_K_S6.46 GiB6,935,348,9924.205mradermacher
I1-Q4_K_S6.46 GiB6,935,349,2484.205mradermacher
Q4_K_M6.80 GiB7,300,776,5444.427NikolayKozloff
Q4_K_M6.80 GiB7,300,794,1124.427mradermacher
I1-Q4_K_M6.80 GiB7,300,794,3684.427mradermacher
I1-Q4_17.04 GiB7,559,579,6484.584mradermacher
Q5_K_S7.67 GiB8,231,978,7524.991mradermacher
I1-Q5_K_S7.67 GiB8,231,979,0084.991mradermacher
Q5_K_M7.87 GiB8,445,052,6725.120mradermacher
I1-Q5_K_M7.87 GiB8,445,052,9285.120mradermacher
Q6_K9.00 GiB9,660,827,3925.858mradermacher
I1-Q6_K9.00 GiB9,660,827,6485.858mradermacher
Q8_011.65 GiB12,510,228,3527.585mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 6.91 GiB. The real file is 6.80 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does translategemma-12b-it need?
Q4_K_M is exactly 7,300,776,544 bytes (6.80 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of translategemma-12b-it should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.