Best local AI models for 8GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 8GB card gives you about 7.44 GiB to work with after driver overhead. 20 indexed models fit at 32K context — the largest being s2-pro at 4.6B parameters in Q8_0.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 8GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Ace-Step1.5speech synthesisF32160M2.71 GiB4.73 GiB
Qwen3-TTS-12Hz-0.6B-Basespeech synthesisQ4_K_M915M5.57 GiB1.87 GiB
OmniVoicespeech synthesisF32613M3.82 GiB3.62 GiB
orpheus-3b-0.1-ftspeech synthesisQ6_K3.8B6.84 GiB0.60 GiB
VieNeu-TTS-0.3Bspeech synthesisQ8_0244M1.94 GiB5.50 GiB
neutts-airspeech synthesisQ8_0748M1.90 GiB5.54 GiB
s2-proMoEspeech synthesisQ8_04.6B6.07 GiB1.37 GiB
Fun-CosyVoice3-0.5B-2512speech synthesisF164.37 GiB3.07 GiB
VoxCPM2speech synthesisF162.3B5.57 GiB1.87 GiB
Kokoro-82Mspeech synthesisF1682M1.00 GiB6.44 GiB
VieNeu-TTSspeech synthesisQ4_0553M1.54 GiB5.90 GiB
csm-1bspeech synthesisF161.6B5.13 GiB2.31 GiB
VibeVoice-Realtime-0.5Bspeech synthesisF161.0B2.74 GiB4.70 GiB
MOSS-TTS-Local-Transformer-v1.5speech synthesisQ8_04.6B6.41 GiB1.03 GiB
Qwen3-TTS-12Hz-1.7B-Basespeech synthesisF161.9B4.44 GiB3.00 GiB
orpheus-3b-0.1-pretrainedspeech synthesisQ6_K_L3.8B7.43 GiB0.01 GiB
Qwen3-TTS-12Hz-1.7B-VoiceDesignspeech synthesisF161.9B4.42 GiB3.02 GiB
VibeVoice-1.5Bspeech synthesisBF162.7B5.89 GiB1.55 GiB
VoxCPM-0.5Bspeech synthesisQ8_0728M4.92 GiB2.52 GiB
Qwen3-TTS-12Hz-1.7B-CustomVoicespeech synthesisF161.9B4.42 GiB3.02 GiB
Spec sheetPredictedwhat these mean

This page models a generic 8GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.