Best models under 4GB
Models whose smallest published quantization fits in under 4GB. These run on laptops, older cards, and integrated graphics.
From the file· live filter over real dataFrom the file· 40 models
How this is ranked
Ranked by downloads among models with a real quantization under 4GB. Size is measured from published file bytes, not estimated.
Best models under 4GB
| # | Model | Params○ | Smallest quant● | Smallest quant● |
|---|---|---|---|---|
| 1 | Qwen3.5-9B | 9.7B | 2.97 GiB | 2.97 GiB |
| 2 | gemma-4-12B-it | 12.0B | 3.92 GiB | 3.92 GiB |
| 3 | nemotron-3.5-asr-streaming-0.6b | 638M | 0.38 GiB | 0.38 GiB |
| 4 | Qwen3.5-4B | 4.7B | 1.42 GiB | 1.42 GiB |
| 5 | parakeet-unified-en-0.6b | 618M | 0.44 GiB | 0.44 GiB |
| 6 | gemma-4-E4B-it | 8.0B | 3.30 GiB | 3.30 GiB |
| 7 | Qwen3-4B | 4.0B | 1.01 GiB | 1.01 GiB |
| 8 | FLUX.2-klein-9B | 9.1B | 3.71 GiB | 3.71 GiB |
| 9 | Qwen3-8B | 8.2B | 2.12 GiB | 2.12 GiB |
| 10 | cohere-transcribe-03-2026 | 2.1B | 1.41 GiB | 1.41 GiB |
| 11 | HyperCLOVAX-SEED-Text-Instruct-1.5B | 1.6B | 1.06 GiB | 1.06 GiB |
| 12 | Qwen3.5-0.8B | 873M | 0.31 GiB | 0.31 GiB |
| 13 | Llama-3.2-1B-Instruct | 1.2B | 0.39 GiB | 0.39 GiB |
| 14 | gemma-4-E2B-it | 5.1B | 2.13 GiB | 2.13 GiB |
| 15 | parakeet-tdt-0.6b-v3 | 627M | 0.39 GiB | 0.39 GiB |
| 16 | Qwen3-VL-8B-Instruct-abliterated-v1 | 8.8B | 1.97 GiB | 1.97 GiB |
| 17 | Llama-3.1-8B-Instruct | 8.0B | 2.02 GiB | 2.02 GiB |
| 18 | ced-base | 86M | 0.12 GiB | 0.12 GiB |
| 19 | Ace-Step1.5 | 160M | 0.04 GiB | 0.04 GiB |
| 20 | Qwen2.5-7B-Instruct | 7.6B | 2.59 GiB | 2.59 GiB |
| 21 | Qwen3-TTS-12Hz-0.6B-Base | 915M | 0.50 GiB | 0.50 GiB |
| 22 | Qwythos-9B-v2 | 9.7B | 3.64 GiB | 3.64 GiB |
| 23 | UI-TARS-1.5-7B | 8.3B | 2.81 GiB | 2.81 GiB |
| 24 | gemma-3-1b-it | 1000M | 0.52 GiB | 0.52 GiB |
| 25 | gemma-4-E2B-it-qat-q4_0-unquantized | 5.1B | 3.12 GiB | 3.12 GiB |
| 26 | Qwen3-1.7B | 2.0B | 0.50 GiB | 0.50 GiB |
| 27 | whisper-medium | 764M | 0.25 GiB | 0.25 GiB |
| 28 | embeddinggemma-300m | 303M | 0.26 GiB | 0.26 GiB |
| 29 | Llama-3.2-3B-Instruct | 3.2B | 0.85 GiB | 0.85 GiB |
| 30 | Qwen3-0.6B | 752M | 0.20 GiB | 0.20 GiB |
| 31 | Qwen3-14B | 14.8B | 3.56 GiB | 3.56 GiB |
| 32 | Voxtral-Mini-4B-Realtime-2602 | 4.4B | 2.35 GiB | 2.35 GiB |
| 33 | Wan2.1-T2V-1.3B | 1.4B | 0.61 GiB | 0.61 GiB |
| 34 | LFM2.5-1.2B-Instruct | 1.2B | 0.45 GiB | 0.45 GiB |
| 35 | Qwen2.5-Coder-7B-Instruct | 7.6B | 2.59 GiB | 2.59 GiB |
| 36 | Qwen3-VL-4B-Instruct | 4.4B | 1.01 GiB | 1.01 GiB |
| 37 | Qwen2.5-1.5B-Instruct | 1.5B | 0.56 GiB | 0.56 GiB |
| 38 | Jan-v3-4B-base-instruct | 4.4B | 1.56 GiB | 1.56 GiB |
| 39 | gemma-3-4b-it | 4.3B | 1.43 GiB | 1.43 GiB |
| 40 | whisper-large-v3 | 1.5B | 0.49 GiB | 0.49 GiB |
Best local models for codingBest local reasoning modelsBest models for an 8GB cardBest models for long contextBest permissively licensed modelsLargest models with a runnable quantizationBest local speech recognition modelsBest local text-to-speech modelsBest local vision-language modelsBest local image generation models