Model comparison

Qwen3.5-REAP-262B-A17B vs DeepSeek-V4-Flash

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 14.9× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

Qwen3.5-REAP-262B-A17BDeepSeek-V4-Flash
Parameters262B291B
Architectureqwen35moedeepseek4
Layers6043
Native context262,1441,048,576
Mixture of expertsyes, 333 expertsyes, 256 experts
Quantizations published1012
Smallest quantization64.50 GiB76.87 GiB
Q4_K_M147.71 GiB
Licenceapache-2.0mit

KV cache by context

the term that decides long-context viability
ContextQwen3.5-REAP-262B-A17BDeepSeek-V4-FlashRatio
4,0960.12 GiB0.06 GiB1.86×
8,1920.23 GiB0.06 GiB3.72×
16,3840.47 GiB0.06 GiB7.44×
32,7680.94 GiB0.06 GiB14.89×
65,5361.88 GiB0.06 GiB29.77×
131,0723.75 GiB0.06 GiB59.54×