Model comparison

GLM-5 vs GLM-5.2

These two publish different quantization sets; the table below has the exact sizes.

From the file· summed bytes, KV per layer

Side by side

GLM-5GLM-5.2
Parameters754B753B
Architectureglm-dsaglm-dsa
Layers7878
Native context202,7521,048,576
Mixture of expertsyes, 256 expertsyes, 256 experts
Quantizations published2120
Smallest quantization164.05 GiB169.33 GiB
Q4_K_M424.57 GiB
Licencemitmit

KV cache by context

the term that decides long-context viability
ContextGLM-5GLM-5.2Ratio
4,0960.34 GiB0.34 GiB1.00×
8,1920.69 GiB0.69 GiB1.00×
16,3841.37 GiB1.37 GiB1.00×
32,7682.74 GiB2.74 GiB1.00×
65,5365.48 GiB5.48 GiB1.00×
131,07210.97 GiB10.97 GiB1.00×