Model comparison

granite-speech-4.1-2b vs cohere-transcribe-03-2026

At Q4_K_M, granite-speech-4.1-2b is the smaller download — 1,139,247,200 bytes against 1,558,162,944.

From the file· summed bytes, KV per layer

Side by side

granite-speech-4.1-2bcohere-transcribe-03-2026
Parameters2.3B2.1B
Architecturegranite_speechcohere_asr
Layers40
Native context4,096
Mixture of expertsnono
Quantizations published1311
Smallest quantization1.06 GiB1.41 GiB
Q4_K_M1.06 GiB1.45 GiB
Licenceapache-2.0apache-2.0

KV cache by context

the term that decides long-context viability
Contextgranite-speech-4.1-2bcohere-transcribe-03-2026Ratio
4,0960.31 GiB
8,1920.63 GiB
16,3841.25 GiB
32,7682.50 GiB
65,5365.00 GiB
131,07210.00 GiB