Best local AI models for 6GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 6GB card gives you about 5.58 GiB to work with after driver overhead. 3 indexed models fit at 32K context — the largest being HunyuanVideo-1.5 at 8.3B parameters in Q4_K_S.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 6GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Wan2.2-TI2V-5Bvideo generationQ6_K5.0B4.76 GiB0.82 GiB
HunyuanVideo-1.5video generationQ4_K_S8.3B5.43 GiB0.15 GiB
Wan2.2-TI2V-5B-Turbovideo generationQ6_K5.0B4.76 GiB0.82 GiB
Spec sheetPredictedwhat these mean

This page models a generic 6GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.