Best local AI models for 48GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 48GB card gives you about 44.64 GiB to work with after driver overhead. 39 indexed models fit at 32K context — the largest being Voxtral-Small-24B-2507 at 24.3B parameters in Q8_0.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 48GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| nemotron-3.5-asr-streaming-0.6b | speech recognition | F32 | 638M | 3.22 GiB | 41.42 GiB |
| parakeet-tdt-0.6b-v3 | speech recognition | F32 | 627M | 3.18 GiB | 41.46 GiB |
| whisper-medium | speech recognition | F32 | 764M | 3.69 GiB | 40.95 GiB |
| Voxtral-Mini-4B-Realtime-2602 | speech recognition | F16 | 4.4B | 12.33 GiB | 32.31 GiB |
| whisper-large-v3 | speech recognition | F16 | 1.5B | 3.74 GiB | 40.90 GiB |
| whisper-large-v3-turbo | speech recognition | F16 | 809M | 2.36 GiB | 42.28 GiB |
| Qwen3-ASR-1.7B | speech recognition | F16 | 2.3B | 5.23 GiB | 39.41 GiB |
| Voxtral-Small-24B-2507 | speech recognition | Q8_0 | 24.3B | 29.96 GiB | 14.68 GiB |
| Qwen3-ASR-0.6B | speech recognition | F16 | 938M | 2.32 GiB | 42.32 GiB |
| GigaAM-v3 | speech recognition | F32 | 223M | 1.67 GiB | 42.97 GiB |
| parakeet-ctc-0.6b | speech recognition | F16 | 609M | 16.97 GiB | 27.67 GiB |
| granite-speech-4.1-2b-nar | speech recognition | F16 | 2.3B | 8.65 GiB | 35.99 GiB |
| whisper-small | speech recognition | F32 | 242M | 1.75 GiB | 42.89 GiB |
| granite-speech-4.1-2b | speech recognition | F16 | 2.3B | 8.48 GiB | 36.16 GiB |
| whisper-large | speech recognition | F32 | 1.5B | 6.60 GiB | 38.04 GiB |
| Voxtral-Mini-3B-2507 | speech recognition | F16 | 4.7B | 13.29 GiB | 31.35 GiB |
| canary-1b-flash | speech recognition | F32 | 811M | 4.16 GiB | 40.48 GiB |
| whisper-large-v2 | speech recognition | F32 | 1.5B | 6.60 GiB | 38.04 GiB |
| canary-qwen-2.5b | speech recognition | F16 | 2.6B | 6.15 GiB | 38.49 GiB |
| granite-speech-4.1-2b-plus | speech recognition | F16 | 2.1B | 8.50 GiB | 36.14 GiB |
| Breeze-ASR-25 | speech recognition | F16 | 1.5B | 3.74 GiB | 40.90 GiB |
| nemotron-speech-streaming-en-0.6b | speech recognition | F32 | 618M | 3.15 GiB | 41.49 GiB |
| granite-4.0-1b-speech | speech recognition | F16 | 2.3B | 7.60 GiB | 37.04 GiB |
| parakeet-ctc-1.1b | speech recognition | F32 | 1.1B | 4.80 GiB | 39.84 GiB |
| parakeet-rnnt-1.1b | speech recognition | F32 | 1.1B | 4.83 GiB | 39.81 GiB |
| whisper-base | speech recognition | F32 | 73M | 1.12 GiB | 43.52 GiB |
| moonshine-streaming-medium | speech recognition | F32 | 266M | 2.85 GiB | 41.79 GiB |
| whisper-medium.en | speech recognition | F32 | 764M | 3.69 GiB | 40.95 GiB |
| parakeet-rnnt-0.6b | speech recognition | F32 | 617M | 3.14 GiB | 41.50 GiB |
| whisper-small.en | speech recognition | F32 | 242M | 1.75 GiB | 42.89 GiB |
| moonshine-streaming-small | speech recognition | F32 | 140M | 1.91 GiB | 42.73 GiB |
| whisper-tiny | speech recognition | F32 | 38M | 0.99 GiB | 43.65 GiB |
| MOSS-Transcribe-Diarize | speech recognition | F32 | 909M | 7.66 GiB | 36.98 GiB |
| GLM-ASR-Nano-2512 | speech recognition | BF16 | 2.3B | 5.51 GiB | 39.13 GiB |
| ARK-ASR-3B | speech recognition | F16 | 4.1B | 8.93 GiB | 35.71 GiB |
| whisper-base.en | speech recognition | F32 | 73M | 1.12 GiB | 43.52 GiB |
| moonshine-streaming-tiny | speech recognition | F32 | 44M | 1.16 GiB | 43.48 GiB |
| moonshine-base | speech recognition | F32 | 62M | 0.99 GiB | 43.65 GiB |
| Qwen3-ForcedAligner-0.6B | speech recognition | F16 | 918M | 2.56 GiB | 42.08 GiB |
This page models a generic 48GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.