Best local AI models for 12GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 12GB card gives you about 11.16 GiB to work with after driver overhead. 21 indexed models fit at 32K context — the largest being s2-pro at 4.6B parameters in F16.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 12GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Ace-Step1.5speech synthesisQ4_K160M9.65 GiB1.51 GiB
Qwen3-TTS-12Hz-0.6B-Basespeech synthesisQ8_0915M8.68 GiB2.48 GiB
OmniVoicespeech synthesisF32613M3.82 GiB7.34 GiB
orpheus-3b-0.1-ftspeech synthesisF163.8B10.47 GiB0.69 GiB
VieNeu-TTS-0.3Bspeech synthesisQ8_0244M1.94 GiB9.22 GiB
neutts-airspeech synthesisQ8_0748M1.90 GiB9.26 GiB
s2-proMoEspeech synthesisF164.6B10.05 GiB1.11 GiB
Fun-CosyVoice3-0.5B-2512speech synthesisF164.37 GiB6.79 GiB
VoxCPM2speech synthesisF162.3B5.57 GiB5.59 GiB
Kokoro-82Mspeech synthesisF1682M1.00 GiB10.16 GiB
VieNeu-TTSspeech synthesisQ4_0553M1.54 GiB9.62 GiB
csm-1bspeech synthesisF161.6B5.13 GiB6.03 GiB
VibeVoice-Realtime-0.5Bspeech synthesisF161.0B2.74 GiB8.42 GiB
tada-3b-mlspeech synthesisQ4_K4.2B10.08 GiB1.08 GiB
MOSS-TTS-Local-Transformer-v1.5speech synthesisF164.6B9.58 GiB1.58 GiB
Qwen3-TTS-12Hz-1.7B-Basespeech synthesisF161.9B4.44 GiB6.72 GiB
orpheus-3b-0.1-pretrainedspeech synthesisQ8_03.8B8.06 GiB3.10 GiB
Qwen3-TTS-12Hz-1.7B-VoiceDesignspeech synthesisF161.9B4.42 GiB6.74 GiB
VibeVoice-1.5Bspeech synthesisF322.7B10.92 GiB0.24 GiB
VoxCPM-0.5Bspeech synthesisQ8_0728M4.92 GiB6.24 GiB
Qwen3-TTS-12Hz-1.7B-CustomVoicespeech synthesisF161.9B4.42 GiB6.74 GiB
Spec sheetPredictedwhat these mean

This page models a generic 12GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.