Best local AI models for 10GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 10GB card gives you about 9.30 GiB to work with after driver overhead. 20 indexed models fit at 32K context — the largest being s2-pro at 4.6B parameters in Q8_0.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 10GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Ace-Step1.5speech synthesisF32160M2.71 GiB6.59 GiB
Qwen3-TTS-12Hz-0.6B-Basespeech synthesisQ8_0915M8.68 GiB0.62 GiB
OmniVoicespeech synthesisF32613M3.82 GiB5.48 GiB
orpheus-3b-0.1-ftspeech synthesisQ8_03.8B7.58 GiB1.72 GiB
VieNeu-TTS-0.3Bspeech synthesisQ8_0244M1.94 GiB7.36 GiB
neutts-airspeech synthesisQ8_0748M1.90 GiB7.40 GiB
s2-proMoEspeech synthesisQ8_04.6B6.07 GiB3.23 GiB
Fun-CosyVoice3-0.5B-2512speech synthesisF164.37 GiB4.93 GiB
VoxCPM2speech synthesisF162.3B5.57 GiB3.73 GiB
Kokoro-82Mspeech synthesisF1682M1.00 GiB8.30 GiB
VieNeu-TTSspeech synthesisQ4_0553M1.54 GiB7.76 GiB
csm-1bspeech synthesisF161.6B5.13 GiB4.17 GiB
VibeVoice-Realtime-0.5Bspeech synthesisF161.0B2.74 GiB6.56 GiB
MOSS-TTS-Local-Transformer-v1.5speech synthesisF324.6B8.76 GiB0.54 GiB
Qwen3-TTS-12Hz-1.7B-Basespeech synthesisF161.9B4.44 GiB4.86 GiB
orpheus-3b-0.1-pretrainedspeech synthesisQ8_03.8B8.06 GiB1.24 GiB
Qwen3-TTS-12Hz-1.7B-VoiceDesignspeech synthesisF161.9B4.42 GiB4.88 GiB
VibeVoice-1.5Bspeech synthesisBF162.7B5.89 GiB3.41 GiB
VoxCPM-0.5Bspeech synthesisQ8_0728M4.92 GiB4.38 GiB
Qwen3-TTS-12Hz-1.7B-CustomVoicespeech synthesisF161.9B4.42 GiB4.88 GiB
Spec sheetPredictedwhat these mean

This page models a generic 10GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.