Model comparison

step-3.5-flash vs DeepSeek-V4-Flash

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 206.9× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

step-3.5-flashDeepSeek-V4-Flash
Parameters199B291B
Architecturestep35deepseek4
Layers4543
Native context262,1441,048,576
Mixture of expertsnoyes, 256 experts
Quantizations published3012
Smallest quantization38.47 GiB76.87 GiB
Q4_K_M111.70 GiB
Licenceapache-2.0mit

KV cache by context

the term that decides long-context viability
Contextstep-3.5-flashDeepSeek-V4-FlashRatio
4,0962.53 GiB0.06 GiB40.19×
8,1924.03 GiB0.06 GiB64.00×
16,3847.03 GiB0.06 GiB111.63×
32,76813.03 GiB0.06 GiB206.88×
65,53625.03 GiB0.06 GiB397.40×
131,07249.03 GiB0.06 GiB778.42×