Model comparison

starcoder2-3b vs Qwen3-VL-8B-Instruct-abliterated-v1

At Q4_K_M, starcoder2-3b is the smaller download — 1,848,976,448 bytes against 5,027,785,792.

From the file· summed bytes, KV per layer

Side by side

starcoder2-3bQwen3-VL-8B-Instruct-abliterated-v1
Parameters3.0B8.8B
Architecturestarcoder2qwen3vl
Layers30
Native context16,384
Mixture of expertsnono
Quantizations published3748
Smallest quantization1.07 GiB1.97 GiB
Q4_K_M1.72 GiB4.68 GiB
Licencebigcode-openrail-mapache-2.0

KV cache by context

the term that decides long-context viability
Contextstarcoder2-3bQwen3-VL-8B-Instruct-abliterated-v1Ratio
4,096
8,192
16,384
32,768
65,536
131,072