Model comparison

gemma-2-2b-it-abliterated vs Llama-3.1-8B-Instruct

At Q4_K_M, gemma-2-2b-it-abliterated is the smaller download — 1,708,582,784 bytes against 4,920,739,168. At long context the gap widens: gemma-2-2b-it-abliterated's KV cache at 32K is 2.2× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

gemma-2-2b-it-abliteratedLlama-3.1-8B-Instruct
Parameters2.6B8.0B
Architecturegemma2llama
Layers2632
Native context8,192131,072
Mixture of expertsnono
Quantizations published2945
Smallest quantization1.15 GiB2.02 GiB
Q4_K_M1.59 GiB4.58 GiB
Licencegemmallama3.1

KV cache by context

the term that decides long-context viability
Contextgemma-2-2b-it-abliteratedLlama-3.1-8B-InstructRatio
4,0960.41 GiB0.50 GiB1.23×
8,1920.63 GiB1.00 GiB1.58×
16,3841.04 GiB2.00 GiB1.92×
32,7681.85 GiB4.00 GiB2.16×
65,5363.48 GiB8.00 GiB2.30×
131,0726.73 GiB16.00 GiB2.38×