gemma-4-E4B-it-qat-q4_0-unquantized alternatives

gemma-4-E4B-it-qat-q4_0-unquantized's smallest published quantization is 4.80 GiB. The models below do the same job with a different trade — less memory, comparable size, or a more permissive licence.

From the file· sizes from summed file bytes

Meaningfully smaller

under 70% of its smallest quantization
ModelParamsSmallestQuantsLicence
Qwen3.5-9B9.7B2.97 GiB34apache-2.0
Qwen3.5-4B4.7B1.42 GiB57apache-2.0
Qwen3.5-0.8B873M0.31 GiB42apache-2.0
gemma-4-E2B-it5.1B2.13 GiB54apache-2.0
Qwen3-VL-4B-Instruct4.4B1.01 GiB26apache-2.0
Qwen2.5-VL-7B-Instruct8.3B1.93 GiB50apache-2.0
Qwen3-VL-2B-Instruct2.1B0.50 GiB46apache-2.0
gemma-3-12b-it12.2B2.85 GiB25gemma

Comparable in size

within ±40%, so a like-for-like swap
ModelParamsSmallestQuantsLicence
gemma-4-12B-it12.0B3.92 GiB59apache-2.0
Qwythos-9B-Claude-Mythos-5-1M9.4B5.38 GiB16apache-2.0
Qwythos-9B-v29.7B3.64 GiB43apache-2.0
Qwen3.5-9B9.7B4.31 GiB13apache-2.0
Mistral-Small-3.2-24B-Instruct-250624.0B5.18 GiB50apache-2.0
gemma-4-E4B-it-ultra-uncensored-heretic8.0B4.10 GiB24apache-2.0
Qwopus3.5-9B-v3.59.7B3.56 GiB22apache-2.0
Qwen3.5-9B-GLM5.1-Distill-v19.7B4.31 GiB9apache-2.0

More permissively licensed

licences that allow commercial use
ModelParamsSmallestQuantsLicence
Qwen3.6-35B-A3BMoE36.0B8.77 GiB61apache-2.0
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP27.8B10.12 GiB34apache-2.0
Qwen3.5-9B9.7B2.97 GiB34apache-2.0
gemma-4-26B-A4B-itMoE26.5B8.99 GiB46apache-2.0
gemma-4-12B-it12.0B3.92 GiB59apache-2.0
Qwen3.5-4B4.7B1.42 GiB57apache-2.0
Qwythos-9B-Claude-Mythos-5-1M9.4B5.38 GiB16apache-2.0
gemma-4-31B-it31.3B7.95 GiB55apache-2.0