Model comparison

granite-speech-4.1-2b-plus vs cohere-transcribe-03-2026

At Q4_K_M, granite-speech-4.1-2b-plus is the smaller download — 1,023,646,816 bytes against 1,558,162,944.

From the file· summed bytes, KV per layer

Side by side

granite-speech-4.1-2b-pluscohere-transcribe-03-2026
Parameters2.1B2.1B
Architecturegranite_speechcohere_asr
Layers40
Native context4,096
Mixture of expertsnono
Quantizations published1311
Smallest quantization0.95 GiB1.41 GiB
Q4_K_M0.95 GiB1.45 GiB
Licenceapache-2.0apache-2.0

KV cache by context

the term that decides long-context viability
Contextgranite-speech-4.1-2b-pluscohere-transcribe-03-2026Ratio
4,0960.31 GiB
8,1920.63 GiB
16,3841.25 GiB
32,7682.50 GiB
65,5365.00 GiB
131,07210.00 GiB