Model comparison

command-a-plus-05-2026-bf16 vs DeepSeek-V4-Flash

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 22.6× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

command-a-plus-05-2026-bf16DeepSeek-V4-Flash
Parameters219B291B
Architecturecohere2moedeepseek4
Layers3243
Native context200,0001,048,576
Mixture of expertsyes, 128 expertsyes, 256 experts
Quantizations published2512
Smallest quantization45.87 GiB76.87 GiB
Q4_K_M125.81 GiB
Licenceapache-2.0mit

KV cache by context

the term that decides long-context viability
Contextcommand-a-plus-05-2026-bf16DeepSeek-V4-FlashRatio
4,0960.50 GiB0.06 GiB7.94×
8,1920.67 GiB0.06 GiB10.67×
16,3840.92 GiB0.06 GiB14.64×
32,7681.42 GiB0.06 GiB22.57×
65,5362.42 GiB0.06 GiB38.45×
131,0724.42 GiB0.06 GiB70.20×