Best local AI models for 64GB VRAM
Ranked by what actually fits at 32K context, computed from real file bytes.
A 64GB card gives you about 59.52 GiB to work with after driver overhead. 26 indexed models fit at 32K context — the largest being Nemotron-3-Embed-8B-BF16 at 8.0B parameters in BF16.
From the file· fit from summed bytesFrom the file· KV per layer
Fits in 64GB at 32K context
largest quantization that fits, per model
| Model | Modality | Best quant | Params○ | Total◐ | Headroom◐ |
|---|---|---|---|---|---|
| jina-embeddings-v5-text-small | embeddings | F16 | 596M | 5.39 GiB | 54.13 GiB |
| embeddinggemma-300m-qat-q8_0-unquantized | embeddings | Q8_0 | 303M | 1.22 GiB | 58.30 GiB |
| KaLM-embedding-multilingual-mini-instruct-v2.5 | embeddings | Q8_0 | 494M | 1.65 GiB | 57.87 GiB |
| nomic-embed-text-v1.5 | embeddings | F32 | 137M | 2.40 GiB | 57.12 GiB |
| jina-embeddings-v5-text-nano | embeddings | F16 | 212M | 2.30 GiB | 57.22 GiB |
| all-MiniLM-L6-v2 | embeddings | F32 | 23M | 1.13 GiB | 58.39 GiB |
| mxbai-embed-xsmall-v1 | embeddings | F32 | 24M | 1.13 GiB | 58.39 GiB |
| nomic-embed-text-v2-moeMoE | embeddings | F32 | 475M | 2.58 GiB | 56.94 GiB |
| bge-m3 | embeddings | Q8_0 | 567M | 4.37 GiB | 55.15 GiB |
| snowflake-arctic-embed-l-v2.0 | embeddings | F32 | 568M | 5.90 GiB | 53.62 GiB |
| jina-reranker-v1-tiny-en | embeddings | F16 | 33M | 1.01 GiB | 58.51 GiB |
| Qwen3.5-9B-DFlash | embeddings | BF16 | 1.3B | 3.47 GiB | 56.05 GiB |
| Qwen3-Embedding-0.6B | embeddings | F16 | 596M | 5.39 GiB | 54.13 GiB |
| gte-small | embeddings | Q8_0 | 33M | 1.36 GiB | 58.16 GiB |
| Qwen3-Embedding-8B | embeddings | F16 | 7.6B | 19.43 GiB | 40.09 GiB |
| nomic-embed-text-v1 | embeddings | F32 | 137M | 2.40 GiB | 57.12 GiB |
| LFM2.5-Embedding-350M | embeddings | F16 | 354M | 1.82 GiB | 57.70 GiB |
| Qwen3-Embedding-4B | embeddings | F16 | 4.0B | 12.81 GiB | 46.71 GiB |
| Nemotron-3-Embed-8B-BF16 | embeddings | BF16 | 8.0B | 19.91 GiB | 39.61 GiB |
| nomic-embed-code | embeddings | F32 | 7.1B | 28.95 GiB | 30.57 GiB |
| LCO-Embedding-Omni-3B-2605 | embeddings | Q8_0 | 4.7B | 5.30 GiB | 54.22 GiB |
| qwen-indic-v1 | embeddings | F16 | 7.6B | 19.43 GiB | 40.09 GiB |
| Octen-Embedding-4B | embeddings | F16 | 4.0B | 12.81 GiB | 46.71 GiB |
| bge-reranker-v2-m3 | embeddings | Q8_0 | 568M | 4.37 GiB | 55.15 GiB |
| LFM2.5-ColBERT-350M | embeddings | F16 | 353M | 1.82 GiB | 57.70 GiB |
| gte-large | embeddings | Q8_0 | 335M | 4.11 GiB | 55.41 GiB |
This page models a generic 64GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.