Best local text-to-speech models
Speech synthesis models you can run yourself. Ranking these alongside transcription models, as a single 'speech' list would, compares two unrelated tasks.
From the file· live filter over real dataFrom the file· 23 models
How this is ranked
Ranked by downloads among synthesis models. Streaming latency matters more than throughput here, and nobody has measured it on consumer hardware — including us.
Best local text-to-speech models
| # | Model | Params○ | Smallest quant● | Smallest● |
|---|---|---|---|---|
| 1 | Ace-Step1.5 | 160M | 0.04 GiB | 0.04 GiB |
| 2 | Qwen3-TTS-12Hz-0.6B-Base | 915M | 0.50 GiB | 0.50 GiB |
| 3 | OmniVoice | 613M | 0.56 GiB | 0.56 GiB |
| 4 | orpheus-3b-0.1-ft | 3.8B | 0.91 GiB | 0.91 GiB |
| 5 | VieNeu-TTS-0.3B | 244M | 0.19 GiB | 0.19 GiB |
| 6 | neutts-air | 748M | 0.43 GiB | 0.43 GiB |
| 7 | s2-proMoE | 4.6B | 2.40 GiB | 2.40 GiB |
| 8 | Fun-CosyVoice3-0.5B-2512 | — | 0.34 GiB | 0.34 GiB |
| 9 | VoxCPM2 | 2.3B | 1.57 GiB | 1.57 GiB |
| 10 | chatterbox | 264M | 0.17 GiB | 0.17 GiB |
| 11 | Kokoro-82M | 82M | 0.13 GiB | 0.13 GiB |
| 12 | VieNeu-TTS | 553M | 0.39 GiB | 0.39 GiB |
| 13 | csm-1b | 1.6B | 1.34 GiB | 1.34 GiB |
| 14 | VibeVoice-Realtime-0.5B | 1.0B | 0.65 GiB | 0.65 GiB |
| 15 | tada-3b-ml | 4.2B | 5.20 GiB | 5.20 GiB |
| 16 | MOSS-TTS-Local-Transformer-v1.5 | 4.6B | 5.56 GiB | 5.56 GiB |
| 17 | Qwen3-TTS-12Hz-1.7B-Base | 1.9B | 0.94 GiB | 0.94 GiB |
| 18 | orpheus-3b-0.1-pretrained | 3.8B | 1.40 GiB | 1.40 GiB |
| 19 | ZONOS2 | — | 4.58 GiB | 4.58 GiB |
| 20 | Qwen3-TTS-12Hz-1.7B-VoiceDesign | 1.9B | 1.90 GiB | 1.90 GiB |
| 21 | VibeVoice-1.5B | 2.7B | 1.76 GiB | 1.76 GiB |
| 22 | VoxCPM-0.5B | 728M | 2.45 GiB | 2.45 GiB |
| 23 | Qwen3-TTS-12Hz-1.7B-CustomVoice | 1.9B | 1.90 GiB | 1.90 GiB |
Best local models for codingBest local reasoning modelsBest models for an 8GB cardBest models for long contextBest permissively licensed modelsBest models under 4GBLargest models with a runnable quantizationBest local speech recognition modelsBest local vision-language modelsBest local image generation models