Model comparison

granite-4.0-h-350m vs gemma-3-270m-it

At Q4_K_M, granite-4.0-h-350m is the smaller download — 222,662,560 bytes against 253,115,168. At long context the gap widens: gemma-3-270m-it's KV cache at 32K is 1.2× smaller, which usually matters more than the difference in weights.

From the file· summed bytes, KV per layer

Side by side

granite-4.0-h-350mgemma-3-270m-it
Parameters340M268M
Architecturegranitehybridgemma3
Layers3218
Native context32,76832,768
Mixture of expertsnono
Quantizations published2939
Smallest quantization0.15 GiB0.17 GiB
Q4_K_M0.21 GiB0.24 GiB
Licenceapache-2.0gemma

KV cache by context

the term that decides long-context viability
Contextgranite-4.0-h-350mgemma-3-270m-itRatio
4,0960.02 GiB0.03 GiB1.68×
8,1920.03 GiB0.04 GiB1.22×
16,3840.06 GiB0.06 GiB1.02×
32,7680.13 GiB0.11 GiB1.15×
65,5360.25 GiB0.20 GiB1.24×
131,0720.50 GiB0.39 GiB1.28×