G4-Moonlight-Dusk-26B-A4B-heretic alternatives

G4-Moonlight-Dusk-26B-A4B-heretic's smallest published quantization is 7.95 GiB. The models below do the same job with a different trade — less memory, comparable size, or a more permissive licence.

From the file· sizes from summed file bytes

Meaningfully smaller

under 70% of its smallest quantization
ModelParamsSmallestQuantsLicence
gemma-4-E4B-it8.0B3.30 GiB36apache-2.0
Qwen3-4B4.0B1.01 GiB29
Qwen3-8B8.2B2.12 GiB51apache-2.0
HyperCLOVAX-SEED-Text-Instruct-1.5B1.6B1.06 GiB1other
Llama-3.2-1B-Instruct1.2B0.39 GiB39
llama-3-youko-8b8.0B5.34 GiB2llama3
Qwen3-VL-8B-Instruct-abliterated-v18.8B1.97 GiB48apache-2.0
Llama-3.1-8B-Instruct8.0B2.02 GiB45llama3.1

Comparable in size

within ±40%, so a like-for-like swap
ModelParamsSmallestQuantsLicence
Qwen3-Coder-30B-A3B-InstructMoE30.5B7.46 GiB46apache-2.0
Qwen3.6-27B27.8B8.74 GiB40apache-2.0
Qwen3.8-27B27.8B8.39 GiB22apache-2.0
gemma-4-12B-it-qat-q4_0-unquantized12.0B6.50 GiB2apache-2.0
Qwen3-30B-A3B-Thinking-2507MoE30.5B7.05 GiB51apache-2.0
Laguna-XS-2.1MoE33.4B8.76 GiB31openmdw-1.1
Qwen-AgentWorld-35B-A3BMoE34.7B10.71 GiB15apache-2.0
gpt-oss-20bMoE21.5B10.68 GiB15apache-2.0

More permissively licensed

licences that allow commercial use
ModelParamsSmallestQuantsLicence
Qwen3-Coder-30B-A3B-InstructMoE30.5B7.46 GiB46apache-2.0
Qwen3.6-27B27.8B8.74 GiB40apache-2.0
Qwen3.8-27B27.8B8.39 GiB22apache-2.0
DeepSeek-V4-FlashMoE291B76.87 GiB12mit
gemma-4-12B-it-qat-q4_0-unquantized12.0B6.50 GiB2apache-2.0
gemma-4-E4B-it8.0B3.30 GiB36apache-2.0
Qwen3-30B-A3B-Thinking-2507MoE30.5B7.05 GiB51apache-2.0
Qwen3-8B8.2B2.12 GiB51apache-2.0