Best local AI models for 4GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 4GB card gives you about 3.72 GiB to work with after driver overhead. 2 indexed models fit at 32K context — the largest being Wan2.2-TI2V-5B at 5.0B parameters in Q4_0.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 4GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| Wan2.2-TI2V-5B | video generation | Q4_0 | 5.0B | 3.66 GiB | 0.06 GiB |
| Wan2.2-TI2V-5B-Turbo | video generation | Q4_0 | 5.0B | 3.66 GiB | 0.06 GiB |
This page models a generic 4GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.