Best local AI models for 10GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 10GB card gives you about 9.30 GiB to work with after driver overhead. 12 indexed models fit at 32K context — the largest being Wan2.1-VACE-14B at 17.3B parameters in Q3_K_S.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 10GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Wan2.2-Animate-14Bvideo generationQ3_K17.3B8.98 GiB0.32 GiB
Wan2.1-I2V-14B-480Pvideo generationQ3_K_M16.4B8.83 GiB0.47 GiB
Bernini-Rvideo generationQ4_K_S14.3B8.99 GiB0.31 GiB
Wan2.2-Distill-Modelsvideo generationQ4_K_S14.3B8.99 GiB0.31 GiB
Wan2.2-TI2V-5Bvideo generationQ8_05.0B5.87 GiB3.43 GiB
Wan2.1-T2V-14Bvideo generationQ4_014.3B9.25 GiB0.05 GiB
Wan2.1-I2V-14B-720Pvideo generationQ3_K_M16.4B8.83 GiB0.47 GiB
Wan2.1-VACE-14Bvideo generationQ3_K_S17.3B8.14 GiB1.16 GiB
Wan2.1-FLF2V-14B-720Pvideo generationQ3_K_M16.4B8.83 GiB0.47 GiB
HunyuanVideo-1.5video generationQ8_08.3B9.22 GiB0.08 GiB
Wan2.2-TI2V-5B-Turbovideo generationQ8_05.0B5.87 GiB3.43 GiB
SkyReels-V2-DF-14B-540Pvideo generationQ4_K_S14.3B8.99 GiB0.31 GiB
Spec sheetPredictedwhat these mean

This page models a generic 10GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.