Best local AI models for 12GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 12GB card gives you about 11.16 GiB to work with after driver overhead. 14 indexed models fit at 32K context — the largest being Wan2.1-VACE-14B at 17.3B parameters in Q4_K_S.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 12GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Wan2.2-Animate-14Bvideo generationQ4_K_S17.3B10.70 GiB0.46 GiB
Wan2.1-I2V-14B-480Pvideo generationQ4_116.4B11.15 GiB0.01 GiB
Bernini-Rvideo generationQ5_114.3B11.10 GiB0.06 GiB
Wan2.2-Distill-Modelsvideo generationQ5_114.3B11.10 GiB0.06 GiB
Wan2.2-TI2V-5Bvideo generationQ8_05.0B5.87 GiB5.29 GiB
Wan2.1-T2V-14Bvideo generationQ5_014.3B10.88 GiB0.28 GiB
Wan2.1-I2V-14B-720Pvideo generationQ4_116.4B11.15 GiB0.01 GiB
Wan2.1-VACE-14Bvideo generationQ4_K_S17.3B10.66 GiB0.50 GiB
Wan2.2-S2V-14Bvideo generationIQ3_XXS16.3B10.82 GiB0.34 GiB
Wan2.1-FLF2V-14B-720Pvideo generationQ4_116.4B11.16 GiB0.00 GiB
JoyAI-Echovideo generationQ5_112.2B10.08 GiB1.08 GiB
HunyuanVideo-1.5video generationQ8_08.3B9.22 GiB1.94 GiB
Wan2.2-TI2V-5B-Turbovideo generationQ8_05.0B5.87 GiB5.29 GiB
SkyReels-V2-DF-14B-540Pvideo generationQ5_114.3B11.10 GiB0.06 GiB
Spec sheetPredictedwhat these mean

This page models a generic 12GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.