Best local AI models for 48GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 48GB card gives you about 44.64 GiB to work with after driver overhead. 2 indexed models fit at 32K context — the largest being HunyuanImage-2.1 at 17.5B parameters in Q3_K_M.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 48GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| HunyuanImage-2.1 | image generation | Q3_K_M | 17.5B | 40.48 GiB | 4.16 GiB |
| Janus-Pro-7B | image generation | F16 | 7.4B | 28.70 GiB | 15.94 GiB |
This page models a generic 48GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.