Model comparison

gemma-2-2b-it-abliterated vs Qwen3-VL-8B-Instruct-abliterated-v1

At Q4_K_M, gemma-2-2b-it-abliterated is the smaller download — 1,708,582,784 bytes against 5,027,785,792.

From the file· summed bytes, KV per layer

Side by side

gemma-2-2b-it-abliteratedQwen3-VL-8B-Instruct-abliterated-v1
Parameters2.6B8.8B
Architecturegemma2qwen3vl
Layers26
Native context8,192
Mixture of expertsnono
Quantizations published2948
Smallest quantization1.15 GiB1.97 GiB
Q4_K_M1.59 GiB4.68 GiB
Licencegemmaapache-2.0

KV cache by context

the term that decides long-context viability
Contextgemma-2-2b-it-abliteratedQwen3-VL-8B-Instruct-abliterated-v1Ratio
4,0960.41 GiB
8,1920.63 GiB
16,3841.04 GiB
32,7681.85 GiB
65,5363.48 GiB
131,0726.73 GiB
gemma-2-2b-it-abliterated vs Qwen3-VL-8B-Instruct-abliterated-v1 — size, memory and hardware fit — ossmodeldb