Best local speech recognition models
Transcription models that run on your own machine. These are among the smallest models in local AI, and most run comfortably without a GPU.
From the file· live filter over real dataFrom the file· 40 models
How this is ranked
Ranked by downloads. We publish sizes but not real-time factors — the speed figures that circulate for these were measured on datacenter hardware and say nothing about a laptop.
Best local speech recognition models
| # | Model | Params○ | Smallest quant● | Smallest● |
|---|---|---|---|---|
| 1 | nemotron-3.5-asr-streaming-0.6b | 638M | 0.38 GiB | 0.38 GiB |
| 2 | parakeet-unified-en-0.6b | 618M | 0.44 GiB | 0.44 GiB |
| 3 | cohere-transcribe-03-2026 | 2.1B | 1.41 GiB | 1.41 GiB |
| 4 | parakeet-tdt-0.6b-v3 | 627M | 0.39 GiB | 0.39 GiB |
| 5 | whisper-medium | 764M | 0.25 GiB | 0.25 GiB |
| 6 | Voxtral-Mini-4B-Realtime-2602 | 4.4B | 2.35 GiB | 2.35 GiB |
| 7 | whisper-large-v3 | 1.5B | 0.49 GiB | 0.49 GiB |
| 8 | canary-180m-flash | 189M | 0.13 GiB | 0.13 GiB |
| 9 | whisper-large-v3-turbo | 809M | 0.27 GiB | 0.27 GiB |
| 10 | Qwen3-ASR-1.7B | 2.3B | 1.23 GiB | 1.23 GiB |
| 11 | Voxtral-Small-24B-2507 | 24.3B | 6.71 GiB | 6.71 GiB |
| 12 | Qwen3-ASR-0.6B | 938M | 0.55 GiB | 0.55 GiB |
| 13 | parakeet-tdt-0.6b-v2 | 618M | 0.37 GiB | 0.37 GiB |
| 14 | GigaAM-v3 | 223M | 0.17 GiB | 0.17 GiB |
| 15 | canary-1b-v2 | 980M | 0.37 GiB | 0.37 GiB |
| 16 | parakeet-ctc-0.6b | 609M | 0.44 GiB | 0.44 GiB |
| 17 | granite-speech-4.1-2b-nar | 2.3B | 1.45 GiB | 1.45 GiB |
| 18 | whisper-small | 242M | 0.08 GiB | 0.08 GiB |
| 19 | granite-speech-4.1-2b | 2.3B | 1.06 GiB | 1.06 GiB |
| 20 | whisper-large | 1.5B | 0.93 GiB | 0.93 GiB |
| 21 | Fun-ASR-MLT-Nano-2512 | 830M | 0.52 GiB | 0.52 GiB |
| 22 | Voxtral-Mini-3B-2507 | 4.7B | 1.45 GiB | 1.45 GiB |
| 23 | canary-1b-flash | 811M | 0.63 GiB | 0.63 GiB |
| 24 | whisper-large-v2 | 1.5B | 0.93 GiB | 0.93 GiB |
| 25 | Fun-ASR-Nano-2512 | 830M | 0.52 GiB | 0.52 GiB |
| 26 | canary-qwen-2.5b | 2.6B | 1.62 GiB | 1.62 GiB |
| 27 | parakeet-tdt-1.1b | 1.1B | 0.64 GiB | 0.64 GiB |
| 28 | SenseVoiceSmall | 234M | 0.13 GiB | 0.13 GiB |
| 29 | granite-speech-4.1-2b-plus | 2.1B | 0.95 GiB | 0.95 GiB |
| 30 | Breeze-ASR-25 | 1.5B | 0.93 GiB | 0.93 GiB |
| 31 | nemotron-speech-streaming-en-0.6b | 618M | 0.44 GiB | 0.44 GiB |
| 32 | granite-4.0-1b-speech | 2.3B | 1.49 GiB | 1.49 GiB |
| 33 | parakeet-ctc-1.1b | 1.1B | 0.76 GiB | 0.76 GiB |
| 34 | parakeet-rnnt-1.1b | 1.1B | 0.64 GiB | 0.64 GiB |
| 35 | canary-1b | 1.0B | 0.68 GiB | 0.68 GiB |
| 36 | whisper-base | 73M | 0.03 GiB | 0.03 GiB |
| 37 | moonshine-streaming-medium | 266M | 0.28 GiB | 0.28 GiB |
| 38 | whisper-medium.en | 764M | 0.47 GiB | 0.47 GiB |
| 39 | parakeet-rnnt-0.6b | 617M | 0.37 GiB | 0.37 GiB |
| 40 | parakeet-tdt_ctc-1.1b | 1.1B | 0.64 GiB | 0.64 GiB |
Best local models for codingBest local reasoning modelsBest models for an 8GB cardBest models for long contextBest permissively licensed modelsBest models under 4GBLargest models with a runnable quantizationBest local text-to-speech modelsBest local vision-language modelsBest local image generation models