Model comparison

Qwen2-VL-2B-Instruct vs Unlimited-OCR

At Q4_K_M, Qwen2-VL-2B-Instruct is the smaller download — 986,046,944 bytes against 1,950,321,792.

From the file· summed bytes, KV per layer

Side by side

Qwen2-VL-2B-InstructUnlimited-OCR
Parameters2.2B3.3B
Architectureqwen2vldeepseek2-ocr
Layers2812
Native context32,76832,768
Mixture of expertsnoyes, 64 experts
Quantizations published2521
Smallest quantization0.56 GiB1.15 GiB
Q4_K_M0.92 GiB1.82 GiB
Licenceapache-2.0mit

KV cache by context

the term that decides long-context viability
ContextQwen2-VL-2B-InstructUnlimited-OCRRatio
4,0960.11 GiB
8,1920.22 GiB
16,3840.44 GiB
32,7680.88 GiB
65,5361.75 GiB
131,0723.50 GiB