Best local AI models for 48GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 48GB card gives you about 44.64 GiB to work with after driver overhead. 26 indexed models fit at 32K context — the largest being Nemotron-3-Embed-8B-BF16 at 8.0B parameters in BF16.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 48GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| jina-embeddings-v5-text-small | embeddings | F16 | 596M | 5.39 GiB | 39.25 GiB |
| embeddinggemma-300m-qat-q8_0-unquantized | embeddings | Q8_0 | 303M | 1.22 GiB | 43.42 GiB |
| KaLM-embedding-multilingual-mini-instruct-v2.5 | embeddings | Q8_0 | 494M | 1.65 GiB | 42.99 GiB |
| nomic-embed-text-v1.5 | embeddings | F32 | 137M | 2.40 GiB | 42.24 GiB |
| jina-embeddings-v5-text-nano | embeddings | F16 | 212M | 2.30 GiB | 42.34 GiB |
| all-MiniLM-L6-v2 | embeddings | F32 | 23M | 1.13 GiB | 43.51 GiB |
| mxbai-embed-xsmall-v1 | embeddings | F32 | 24M | 1.13 GiB | 43.51 GiB |
| nomic-embed-text-v2-moeMoE | embeddings | F32 | 475M | 2.58 GiB | 42.06 GiB |
| bge-m3 | embeddings | Q8_0 | 567M | 4.37 GiB | 40.27 GiB |
| snowflake-arctic-embed-l-v2.0 | embeddings | F32 | 568M | 5.90 GiB | 38.74 GiB |
| jina-reranker-v1-tiny-en | embeddings | F16 | 33M | 1.01 GiB | 43.63 GiB |
| Qwen3.5-9B-DFlash | embeddings | BF16 | 1.3B | 3.47 GiB | 41.17 GiB |
| Qwen3-Embedding-0.6B | embeddings | F16 | 596M | 5.39 GiB | 39.25 GiB |
| gte-small | embeddings | Q8_0 | 33M | 1.36 GiB | 43.28 GiB |
| Qwen3-Embedding-8B | embeddings | F16 | 7.6B | 19.43 GiB | 25.21 GiB |
| nomic-embed-text-v1 | embeddings | F32 | 137M | 2.40 GiB | 42.24 GiB |
| LFM2.5-Embedding-350M | embeddings | F16 | 354M | 1.82 GiB | 42.82 GiB |
| Qwen3-Embedding-4B | embeddings | F16 | 4.0B | 12.81 GiB | 31.83 GiB |
| Nemotron-3-Embed-8B-BF16 | embeddings | BF16 | 8.0B | 19.91 GiB | 24.73 GiB |
| nomic-embed-code | embeddings | F32 | 7.1B | 28.95 GiB | 15.69 GiB |
| LCO-Embedding-Omni-3B-2605 | embeddings | Q8_0 | 4.7B | 5.30 GiB | 39.34 GiB |
| qwen-indic-v1 | embeddings | F16 | 7.6B | 19.43 GiB | 25.21 GiB |
| Octen-Embedding-4B | embeddings | F16 | 4.0B | 12.81 GiB | 31.83 GiB |
| bge-reranker-v2-m3 | embeddings | Q8_0 | 568M | 4.37 GiB | 40.27 GiB |
| LFM2.5-ColBERT-350M | embeddings | F16 | 353M | 1.82 GiB | 42.82 GiB |
| gte-large | embeddings | Q8_0 | 335M | 4.11 GiB | 40.53 GiB |
This page models a generic 48GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.