Model comparison

llm-surgery-dark-arts-gpt-oss-60b-96a12 vs DeepSeek-V4-Flash

These two publish different quantization sets; the table below has the exact sizes. At long context the gap widens: DeepSeek-V4-Flash's KV cache at 32K is 12.2× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

llm-surgery-dark-arts-gpt-oss-60b-96a12DeepSeek-V4-Flash
Parameters60.9B291B
Architecturegpt-ossdeepseek4
Layers2443
Native context131,0721,048,576
Mixture of expertsyes, 96 expertsyes, 256 experts
Quantizations published3412
Smallest quantization31.28 GiB76.87 GiB
Q4_K_M41.48 GiB
Licenceapache-2.0mit

KV cache by context

the term that decides long-context viability
Contextllm-surgery-dark-arts-gpt-oss-60b-96a12DeepSeek-V4-FlashRatio
4,0960.11 GiB0.06 GiB1.77×
8,1920.21 GiB0.06 GiB3.26×
16,3840.39 GiB0.06 GiB6.23×
32,7680.77 GiB0.06 GiB12.19×
65,5361.52 GiB0.06 GiB24.09×
131,0723.02 GiB0.06 GiB47.91×