Model comparison

Qwen3-235B-A22B vs DeepSeek-V4-Flash

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 93.3× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

Qwen3-235B-A22BDeepSeek-V4-Flash
Parameters235B291B
Architectureqwen3moedeepseek4
Layers9443
Native context40,9601,048,576
Mixture of expertsyes, 128 expertsyes, 256 experts
Quantizations published3512
Smallest quantization79.81 GiB76.87 GiB
Q4_K_M132.39 GiB
Licenceapache-2.0mit

KV cache by context

the term that decides long-context viability
ContextQwen3-235B-A22BDeepSeek-V4-FlashRatio
4,0960.73 GiB0.06 GiB11.66×
8,1921.47 GiB0.06 GiB23.32×
16,3842.94 GiB0.06 GiB46.64×
32,7685.88 GiB0.06 GiB93.27×
65,53611.75 GiB0.06 GiB186.54×
131,07223.50 GiB0.06 GiB373.09×