GestaltLabs · vision language

Qwen3.6-35B-A3B-NSC-ACE-SABER

GestaltLabs/Qwen3.6-35B-A3B-NSC-ACE-SABER

Qwen3.6-35B-A3B-NSC-ACE-SABER at Q4_K_M is exactly 21,166,757,664 bytes (19.71 GiB / 21.17 GB) — an effective 4.769 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
35.5B
Architecture
qwen35moe
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K12.05 GiB12,939,593,5042.916GestaltLabs
Q2_K12.34 GiB13,252,289,7922.986slevinw
Q2_K12.34 GiB13,252,289,7922.986GestaltLabs
Q3_K_S14.14 GiB15,182,184,2243.421GestaltLabs
Q3_K_S14.48 GiB15,551,290,6243.504slevinw
Q3_K_S14.48 GiB15,551,290,6243.504GestaltLabs
Q3_K_M15.61 GiB16,764,763,9363.777GestaltLabs
Q3_K_M15.99 GiB17,170,914,5603.869GestaltLabs
Q3_K_M15.99 GiB17,170,914,5603.869slevinw
Q3_K_L16.87 GiB18,115,329,8244.082GestaltLabs
Q3_K_L17.28 GiB18,556,345,6004.181GestaltLabs
Q3_K_L17.28 GiB18,556,345,6004.181slevinw
Q4_K_S18.52 GiB19,889,903,3924.482GestaltLabs
Q4_K_S18.97 GiB20,370,003,2004.590GestaltLabs
Q4_K_S18.97 GiB20,370,003,2004.590slevinw
Q4_K_M19.71 GiB21,166,757,6644.769GestaltLabs
Q4_K_M20.23 GiB21,716,604,1604.893slevinw
Q4_K_M20.23 GiB21,716,604,1604.893GestaltLabs
Q5_K_S22.33 GiB23,981,283,1045.403GestaltLabs
Q5_K_S22.88 GiB24,565,847,2965.535GestaltLabs
Q5_K_S22.88 GiB24,565,847,2965.535slevinw
Q5_K_M23.03 GiB24,729,130,7845.572GestaltLabs
Q5_K_M23.61 GiB25,349,625,0885.712slevinw
Q5_K_M23.61 GiB25,349,625,0885.712GestaltLabs
Q6_K26.56 GiB28,514,152,2246.425GestaltLabs
Q6_K27.20 GiB29,209,709,8246.582slevinw
Q6_K27.20 GiB29,209,709,8246.582GestaltLabs
Q8_034.37 GiB36,903,139,1048.315GestaltLabs
Q8_035.21 GiB37,801,096,4488.517GestaltLabs
Q8_035.21 GiB37,801,096,4488.517slevinw
F1664.61 GiB69,376,637,02415.632GestaltLabs
F1666.19 GiB71,065,941,60016.012GestaltLabs
F1666.19 GiB71,065,941,60016.012slevinw

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 18.60 GiB. The real file is 19.71 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Qwen3.6-35B-A3B-NSC-ACE-SABER need?
Q4_K_M is exactly 21,166,757,664 bytes (19.71 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Qwen3.6-35B-A3B-NSC-ACE-SABER should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.