Best local AI models for 12GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 12GB card gives you about 11.16 GiB to work with after driver overhead. 21 indexed models fit at 32K context — the largest being s2-pro at 4.6B parameters in F16.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 12GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| Ace-Step1.5 | speech synthesis | Q4_K | 160M | 9.65 GiB | 1.51 GiB |
| Qwen3-TTS-12Hz-0.6B-Base | speech synthesis | Q8_0 | 915M | 8.68 GiB | 2.48 GiB |
| OmniVoice | speech synthesis | F32 | 613M | 3.82 GiB | 7.34 GiB |
| orpheus-3b-0.1-ft | speech synthesis | F16 | 3.8B | 10.47 GiB | 0.69 GiB |
| VieNeu-TTS-0.3B | speech synthesis | Q8_0 | 244M | 1.94 GiB | 9.22 GiB |
| neutts-air | speech synthesis | Q8_0 | 748M | 1.90 GiB | 9.26 GiB |
| s2-proMoE | speech synthesis | F16 | 4.6B | 10.05 GiB | 1.11 GiB |
| Fun-CosyVoice3-0.5B-2512 | speech synthesis | F16 | — | 4.37 GiB | 6.79 GiB |
| VoxCPM2 | speech synthesis | F16 | 2.3B | 5.57 GiB | 5.59 GiB |
| Kokoro-82M | speech synthesis | F16 | 82M | 1.00 GiB | 10.16 GiB |
| VieNeu-TTS | speech synthesis | Q4_0 | 553M | 1.54 GiB | 9.62 GiB |
| csm-1b | speech synthesis | F16 | 1.6B | 5.13 GiB | 6.03 GiB |
| VibeVoice-Realtime-0.5B | speech synthesis | F16 | 1.0B | 2.74 GiB | 8.42 GiB |
| tada-3b-ml | speech synthesis | Q4_K | 4.2B | 10.08 GiB | 1.08 GiB |
| MOSS-TTS-Local-Transformer-v1.5 | speech synthesis | F16 | 4.6B | 9.58 GiB | 1.58 GiB |
| Qwen3-TTS-12Hz-1.7B-Base | speech synthesis | F16 | 1.9B | 4.44 GiB | 6.72 GiB |
| orpheus-3b-0.1-pretrained | speech synthesis | Q8_0 | 3.8B | 8.06 GiB | 3.10 GiB |
| Qwen3-TTS-12Hz-1.7B-VoiceDesign | speech synthesis | F16 | 1.9B | 4.42 GiB | 6.74 GiB |
| VibeVoice-1.5B | speech synthesis | F32 | 2.7B | 10.92 GiB | 0.24 GiB |
| VoxCPM-0.5B | speech synthesis | Q8_0 | 728M | 4.92 GiB | 6.24 GiB |
| Qwen3-TTS-12Hz-1.7B-CustomVoice | speech synthesis | F16 | 1.9B | 4.42 GiB | 6.74 GiB |
This page models a generic 12GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.