Best local AI models for 4GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 4GB card gives you about 3.72 GiB to work with after driver overhead. 13 indexed models fit at 32K context — the largest being Qwen3.5-9B-DFlash at 1.3B parameters in BF16.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 4GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| embeddinggemma-300m-qat-q8_0-unquantized | embeddings | Q8_0 | 303M | 1.22 GiB | 2.50 GiB |
| KaLM-embedding-multilingual-mini-instruct-v2.5 | embeddings | Q8_0 | 494M | 1.65 GiB | 2.07 GiB |
| nomic-embed-text-v1.5 | embeddings | F32 | 137M | 2.40 GiB | 1.32 GiB |
| jina-embeddings-v5-text-nano | embeddings | F16 | 212M | 2.30 GiB | 1.42 GiB |
| all-MiniLM-L6-v2 | embeddings | F32 | 23M | 1.13 GiB | 2.59 GiB |
| mxbai-embed-xsmall-v1 | embeddings | F32 | 24M | 1.13 GiB | 2.59 GiB |
| nomic-embed-text-v2-moeMoE | embeddings | F32 | 475M | 2.58 GiB | 1.14 GiB |
| jina-reranker-v1-tiny-en | embeddings | F16 | 33M | 1.01 GiB | 2.71 GiB |
| Qwen3.5-9B-DFlash | embeddings | BF16 | 1.3B | 3.47 GiB | 0.25 GiB |
| gte-small | embeddings | Q8_0 | 33M | 1.36 GiB | 2.36 GiB |
| nomic-embed-text-v1 | embeddings | F32 | 137M | 2.40 GiB | 1.32 GiB |
| LFM2.5-Embedding-350M | embeddings | F16 | 354M | 1.82 GiB | 1.90 GiB |
| LFM2.5-ColBERT-350M | embeddings | F16 | 353M | 1.82 GiB | 1.90 GiB |
This page models a generic 4GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.