Will it fit?

Every quantization of a model against your card, at any context length. Weights are the summed bytes of the real published files; the cache is computed layer by layer from the architecture. Nothing here multiplies a parameter count by a constant.

Context
KV cache

Qwen3.6-35B-A3B fits a GeForce RTX 4090 at UD-Q4_K_S21.35 GiB of 22.32 GiB usable at 32K context. Expect roughly 155 tokens/sec (modeled, ±37%).

QuantWeightsExact bytesKVTotalHeadroomtok/s
BF1666.19 GiB71,065,942,5600.63 GiB67.61 GiB-45.29 GiB
Q8_035.21 GiB37,801,097,5040.63 GiB36.63 GiB-14.31 GiB
UD-Q6_K27.95 GiB30,011,242,7840.63 GiB29.38 GiB-7.06 GiB
Q6_K26.56 GiB28,514,152,2880.63 GiB27.99 GiB-5.67 GiB
UD-Q5_K_M25.23 GiB27,087,812,8960.63 GiB26.66 GiB-4.34 GiB
UD-Q5_K_S23.78 GiB25,538,017,5680.63 GiB25.21 GiB-2.89 GiB
UD-Q4_K_M21.11 GiB22,663,387,4240.63 GiB22.54 GiB-0.22 GiB
UD-Q4_K_S19.92 GiB21,388,319,0080.63 GiB21.35 GiB0.97 GiB155±37%
Q4_K_M19.71 GiB21,166,757,7280.63 GiB21.14 GiB1.18 GiB156±37%
UD-IQ4_NL17.26 GiB18,536,192,2880.63 GiB18.69 GiB3.63 GiB170±37%
UD-IQ4_XS16.96 GiB18,209,036,5760.63 GiB18.39 GiB3.93 GiB172±37%
UD-Q3_K_M15.93 GiB17,104,402,7200.63 GiB17.36 GiB4.96 GiB179±37%
UD-Q3_K_S14.30 GiB15,359,196,1280.63 GiB15.73 GiB6.59 GiB191±37%
UD-IQ3_S14.29 GiB15,346,432,2880.63 GiB15.72 GiB6.60 GiB191±37%
UD-IQ3_XXS13.10 GiB14,069,266,7200.63 GiB14.53 GiB7.79 GiB200±37%
UD-IQ2_M11.07 GiB11,882,969,3760.63 GiB12.50 GiB9.82 GiB219±37%
UD-IQ2_XXS11.01 GiB11,819,399,4560.63 GiB12.44 GiB9.88 GiB220±37%
UD-IQ1_M10.59 GiB11,366,414,6240.63 GiB12.02 GiB10.30 GiB224±37%
From the filePredictedwhat these mean

Weights are exact. The cache is computed per layer. The remaining term — the runtime's working buffers — is modeled, which is why a total carries a band when its parts do not. Speed is a bandwidth roofline, and the ± figure is our measured error against real benchmark runs, published on the methodology page.