Best local AI models for 8GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 8GB card gives you about 7.44 GiB to work with after driver overhead. 8 indexed models fit at 32K context — the largest being Wan2.2-Animate-14B at 17.3B parameters in Q2_K.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 8GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Wan2.2-Animate-14Bvideo generationQ2_K17.3B7.20 GiB0.24 GiB
Bernini-Rvideo generationQ3_K_S14.3B6.91 GiB0.53 GiB
Wan2.2-Distill-Modelsvideo generationQ3_K_S14.3B6.91 GiB0.53 GiB
Wan2.2-TI2V-5Bvideo generationQ8_05.0B5.87 GiB1.57 GiB
Wan2.1-T2V-14Bvideo generationQ3_K_S14.3B7.34 GiB0.10 GiB
HunyuanVideo-1.5video generationQ6_K8.3B7.39 GiB0.05 GiB
Wan2.2-TI2V-5B-Turbovideo generationQ8_05.0B5.87 GiB1.57 GiB
SkyReels-V2-DF-14B-540Pvideo generationQ3_K_S14.3B6.91 GiB0.53 GiB
Spec sheetPredictedwhat these mean

This page models a generic 8GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.