Best local AI models for 6GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 6GB card gives you about 5.58 GiB to work with after driver overhead. 18 indexed models fit at 32K context — the largest being s2-pro at 4.6B parameters in Q6_K.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 6GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Ace-Step1.5speech synthesisF32160M2.71 GiB2.87 GiB
Qwen3-TTS-12Hz-0.6B-Basespeech synthesisQ4_K_M915M5.57 GiB0.01 GiB
OmniVoicespeech synthesisF32613M3.82 GiB1.76 GiB
orpheus-3b-0.1-ftspeech synthesisUD-IQ2_M3.8B5.54 GiB0.04 GiB
VieNeu-TTS-0.3Bspeech synthesisQ8_0244M1.94 GiB3.64 GiB
neutts-airspeech synthesisQ8_0748M1.90 GiB3.68 GiB
s2-proMoEspeech synthesisQ6_K4.6B5.04 GiB0.54 GiB
Fun-CosyVoice3-0.5B-2512speech synthesisF164.37 GiB1.21 GiB
VoxCPM2speech synthesisF162.3B5.57 GiB0.01 GiB
Kokoro-82Mspeech synthesisF1682M1.00 GiB4.58 GiB
VieNeu-TTSspeech synthesisQ4_0553M1.54 GiB4.04 GiB
csm-1bspeech synthesisF161.6B5.13 GiB0.45 GiB
VibeVoice-Realtime-0.5Bspeech synthesisF161.0B2.74 GiB2.84 GiB
Qwen3-TTS-12Hz-1.7B-Basespeech synthesisF161.9B4.44 GiB1.14 GiB
Qwen3-TTS-12Hz-1.7B-VoiceDesignspeech synthesisF161.9B4.42 GiB1.16 GiB
VibeVoice-1.5Bspeech synthesisQ8_02.7B4.75 GiB0.83 GiB
VoxCPM-0.5Bspeech synthesisQ8_0728M4.92 GiB0.66 GiB
Qwen3-TTS-12Hz-1.7B-CustomVoicespeech synthesisF161.9B4.42 GiB1.16 GiB
Spec sheetPredictedwhat these mean

This page models a generic 6GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.