Cannae-AI · text

Gemini-3-Pro-Qwen3.5-35B-A3B

Cannae-AI/Gemini-3-Pro-Qwen3.5-35B-A3B

Gemini-3-Pro-Qwen3.5-35B-A3B at I1-IQ1_S is exactly 7,484,142,400 bytes (6.97 GiB / 7.48 GB) — an effective 1.727 bits per weight, not the nominal 1.

From the file· summed from 1 file(s)
Parameters
34.7B
Architecture
qwen35moe
Context
native (config.json)
License
mit

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S6.97 GiB7,484,142,4001.727mradermacher
I1-IQ1_M7.67 GiB8,239,209,2801.902mradermacher
I1-IQ2_XXS8.85 GiB9,497,654,0802.192mradermacher
I1-IQ2_XS9.79 GiB10,507,031,3602.425mradermacher
I1-IQ2_S9.92 GiB10,652,480,3202.459mradermacher
I1-IQ2_M10.86 GiB11,659,236,1602.691mradermacher
I1-Q2_K_S11.32 GiB12,152,097,6002.805mradermacher
I1-Q2_K12.05 GiB12,939,594,5602.987mradermacher
I1-IQ3_XXS12.69 GiB13,623,759,6803.144mradermacher
I1-IQ3_XS13.49 GiB14,484,144,9603.343mradermacher
I1-Q3_K_S14.14 GiB15,182,185,2803.504mradermacher
I1-IQ3_S14.20 GiB15,250,424,6403.520mradermacher
I1-IQ3_M14.38 GiB15,440,520,0003.564mradermacher
I1-Q3_K_M15.61 GiB16,764,764,9923.869mradermacher
I1-Q3_K_L16.87 GiB18,115,330,8804.181mradermacher
I1-IQ4_XS17.44 GiB18,728,778,5604.323mradermacher
I1-Q4_018.44 GiB19,799,268,1604.570mradermacher
I1-Q4_K_S18.52 GiB19,889,904,4484.591mradermacher
I1-Q4_K_M19.71 GiB21,166,758,7204.886mradermacher
I1-Q4_120.35 GiB21,848,169,2805.043mradermacher
I1-Q5_K_S22.33 GiB23,981,284,1605.535mradermacher
I1-Q5_K_M23.03 GiB24,729,131,8405.708mradermacher
I1-Q6_K26.56 GiB28,514,153,2806.581mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts I1-IQ1_S at roughly 18.16 GiB. The real file is 6.97 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Gemini-3-Pro-Qwen3.5-35B-A3B need?
I1-IQ1_S is exactly 7,484,142,400 bytes (6.97 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Gemini-3-Pro-Qwen3.5-35B-A3B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.