Babett · text

elios_offline

Babett/elios_offline

elios_offline at Q4_K_M is exactly 18,687,058,176 bytes (17.40 GiB / 18.69 GB) — an effective 4.870 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
30.7B
Architecture
gemma4
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S6.66 GiB7,156,484,5761.865mradermacher
I1-IQ1_M7.20 GiB7,725,867,4882.013mradermacher
I1-IQ2_XXS8.08 GiB8,674,839,0082.261mradermacher
I1-IQ2_XS8.88 GiB9,530,354,1442.484mradermacher
I1-IQ2_S9.46 GiB10,157,883,8722.647mradermacher
I1-IQ2_M10.17 GiB10,917,061,0882.845mradermacher
I1-Q2_K_S10.22 GiB10,976,584,1602.861mradermacher
Q2_K11.10 GiB11,916,308,7363.106mradermacher
I1-Q2_K11.10 GiB11,916,308,9603.106mradermacher
I1-IQ3_XXS11.25 GiB12,077,502,9443.147mradermacher
I1-IQ3_XS12.17 GiB13,072,364,0003.407mradermacher
Q3_K_S12.82 GiB13,761,351,9363.586mradermacher
I1-Q3_K_S12.82 GiB13,761,352,1603.586mradermacher
I1-IQ3_S12.82 GiB13,761,352,1603.586mradermacher
I1-IQ3_M13.43 GiB14,424,492,5123.759mradermacher
Q3_K_M14.24 GiB15,287,103,7443.984mradermacher
I1-Q3_K_M14.24 GiB15,287,103,9683.984mradermacher
Q3_K_L15.49 GiB16,628,265,2164.333mradermacher
I1-Q3_K_L15.49 GiB16,628,265,4404.333mradermacher
I1-IQ4_XS15.59 GiB16,735,785,4404.362mradermacher
IQ4_XS15.70 GiB16,862,228,7364.394mradermacher
I1-Q4_016.49 GiB17,701,573,0884.613mradermacher
Q4_K_S16.54 GiB17,763,160,3204.629mradermacher
I1-Q4_K_S16.54 GiB17,763,160,5444.629mradermacher
Q4_K_M17.40 GiB18,687,058,1764.870mradermacher
I1-Q4_K_M17.40 GiB18,687,058,4004.870mradermacher
I1-Q4_118.14 GiB19,481,416,1605.077mradermacher
Q5_K_S19.85 GiB21,311,836,4165.554mradermacher
I1-Q5_K_S19.85 GiB21,311,836,6405.554mradermacher
Q5_K_M20.35 GiB21,845,565,6965.693mradermacher
I1-Q5_K_M20.35 GiB21,845,565,9205.693mradermacher
Q6_K23.47 GiB25,201,479,9366.568mradermacher
I1-Q6_K23.47 GiB25,201,480,1606.568mradermacher
Q8_030.39 GiB32,635,670,7848.505mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 16.08 GiB. The real file is 17.40 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does elios_offline need?
Q4_K_M is exactly 18,687,058,176 bytes (17.40 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of elios_offline should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.