Best local AI models for 4GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
Nothing in the current index fits 4GB at 32K context with f16 KV.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 4GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|
This page models a generic 4GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.