Model comparison

orpheus-3b-0.1-pretrained vs s2-pro

At Q4_K_M, orpheus-3b-0.1-pretrained is the smaller download — 2,363,760,640 bytes against 3,566,165,088.

From the file· summed bytes, KV per layer

Side by side

orpheus-3b-0.1-pretraineds2-pro
Parameters3.8B4.6B
Architecturellamafish-speech
Layers2836
Native context131,072
Mixture of expertsnoyes, 1 experts
Quantizations published457
Smallest quantization1.40 GiB2.40 GiB
Q4_K_M2.20 GiB3.32 GiB
Licenceapache-2.0other

KV cache by context

the term that decides long-context viability
Contextorpheus-3b-0.1-pretraineds2-proRatio
4,0960.44 GiB
8,1920.88 GiB
16,3841.75 GiB
32,7683.50 GiB
65,5367.00 GiB
131,07214.00 GiB