Model comparison

Qwen3-Embedding-0.6B vs nomic-embed-text-v2-moe

At Q4_K_M, nomic-embed-text-v2-moe is the smaller download — 344,120,288 bytes against 396,474,560.

From the file· summed bytes, KV per layer

Side by side

Qwen3-Embedding-0.6Bnomic-embed-text-v2-moe
Parameters596M475M
Architectureqwen3nomic-bert-moe
Layers2812
Native context32,768
Mixture of expertsnoyes, 8 experts
Quantizations published2220
Smallest quantization0.28 GiB0.25 GiB
Q4_K_M0.37 GiB0.32 GiB
Licenceapache-2.0apache-2.0

KV cache by context

the term that decides long-context viability
ContextQwen3-Embedding-0.6Bnomic-embed-text-v2-moeRatio
4,0960.44 GiB
8,1920.88 GiB
16,3841.75 GiB
32,7683.50 GiB
65,5367.00 GiB
131,07214.00 GiB