NVIDIA · consumer
Titan Xp
Titan Xp has 12 GB of VRAM at 548 GB/s — about 11.16 GiB usable after driver and compositor overhead. 1714 of 2118 indexed models fit at 16K context with q8_0 KV.
Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
12 GB
GDDR5X
Bandwidth
548 GB/s
384-bit bus
Tensor FP16
—
dense
TDP
250 W
$1200 MSRP
text 1468video 14vision language 144audio asr 39audio tts 21image 2embedding 26
What fits at 16K context
largest quantization that fits, per model · 1714 of 2118 indexed
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| DeepSeek-V2-Lite-ChatMoE | Q5_0 | 15.7B | 10.10 GiB | 0.25 GiB | 11.16 GiB | 0.00 GiB | 120±37% |
| gemma-4-19B-A4B-it-INSTRUCT-Heretic-UncensoredMoE | I1-IQ4_NL | 19.0B | 9.88 GiB | 0.49 GiB | 11.16 GiB | 0.00 GiB | 38±12.9% |
| gemma-4-19B-A4B-it-The-DECKARD-Heretic-Uncensored-ThinkingMoE | I1-IQ4_NL | 19.0B | 9.88 GiB | 0.49 GiB | 11.16 GiB | 0.00 GiB | 38±12.9% |
| gemma-4-19b-a4b-it-REAP-hereticMoE | I1-IQ4_NL | 19.0B | 9.88 GiB | 0.49 GiB | 11.16 GiB | 0.00 GiB | 38±12.9% |
| Gemma-4-19BMoE | I1-IQ4_NL | 19.0B | 9.88 GiB | 0.49 GiB | 11.16 GiB | 0.00 GiB | 38±12.9% |
| Le-Chaton-Slim-23BMoE | I1-Q3_K_S | 23.3B | 9.48 GiB | 0.86 GiB | 11.16 GiB | 0.00 GiB | 68±37% |
| Wan2.1-FLF2V-14B-720P | Q4_1 | 16.4B | 10.32 GiB | 0.00 GiB | 11.16 GiB | 0.00 GiB | 38±12.9% |
| Wan2.1-I2V-14B-480P | Q4_1 | 16.4B | 10.32 GiB | 0.00 GiB | 11.15 GiB | 0.01 GiB | 38±12.9% |
| Wan2.1-I2V-14B-720P | Q4_1 | 16.4B | 10.32 GiB | 0.00 GiB | 11.15 GiB | 0.01 GiB | 38±12.9% |
| Phi-4-reasoning | Q4_1 | 14.7B | 8.63 GiB | 1.66 GiB | 11.15 GiB | 0.01 GiB | 38±12.9% |
| Phi-4-reasoning-plus | Q4_1 | 14.7B | 8.63 GiB | 1.66 GiB | 11.15 GiB | 0.01 GiB | 38±12.9% |
| phi-4 | Q4_1 | 14.7B | 8.63 GiB | 1.66 GiB | 11.15 GiB | 0.01 GiB | 38±12.9% |
| Ministral-3-14B-Instruct-2512 | Q5_K_M | 13.9B | 8.96 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Ministral-3-14B-Reasoning-2512 | Q5_K_M | 13.9B | 8.96 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Ministral-3-14B-Instruct-2512-BF16-abliterated | I1-Q5_K_M | 13.9B | 8.96 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Ministral-3-14B-abliterated | Q5_K_M | 13.9B | 8.96 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Ministral-3-14B-Instruct-2512-BF16 | Q5_K_M | 13.9B | 8.96 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Ministral-3-14B-Reasoning-2512-Uncensored | I1-Q5_K_M | 13.9B | 8.96 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| glm-4-9b-chat-1m | IQ4_XS | 9.5B | 4.98 GiB | 5.31 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| diffusiongemma-26B-A4B-it-HERETIC-UncensoredMoE | Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| diffusiongemma-26B-A4B-itMoE | Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Voxtral-Small-24B-2507 | Q2_K_L | 24.3B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Devstral-Small-2-24B-Instruct-2512 | Q2_K_L | 24.0B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Dolphin3.0-R1-Mistral-24B | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Dolphin3.0-Mistral-24B | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Cydonia_Vistral | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Dans-PersonalityEngine-V1.2.0-24b | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Dans-PersonalityEngine-V1.3.0-24b | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Devstral-Small-2505 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Mistral-Small-3.2-24B-Instruct-2506 | Q2_K_L | 24.0B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| MS3.2-PaintedFantasy-v3-24B | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Precog-24B-v1 | Q2_K_L | — | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Magidonia-24B-v4.3 | Q2_K_L | — | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Magidonia-24B-v4.2.0 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| MS-2501-DPE-QwQify-v0.1-24B | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| sarvam-m | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Magistral-Small-2506 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Cydonia-24B-v4.1 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Cydonia-24B-v4 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Mistral-Small-3.1-24B-Instruct-2503 | Q2_K_L | 24.0B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Cydonia-24B-v4.3 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Cydonia-24B-v4.2.0 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Mistral-Small-24B-Instruct-2501-abliterated | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Dolphin-Mistral-24B-Venice-Edition | Q2_K_L | 24.0B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Mistral-Small-24B-Instruct-2501 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Hearthfire-24B | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Mistral-Small-24B-ArliAI-RPMax-v1.4 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Codex-24B-Small-3.2 | Q2_K_L | 23.6B | 8.89 GiB | 1.33 GiB | 11.14 GiB | 0.02 GiB | 38±12.9% |
| Qwopus3.6-27B-Coder | IQ2_M | 27.8B | 9.74 GiB | 0.53 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| Goetia-26B-A4B-v1.3-Absolute-Heretic-ARAMoE | I1-Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| Frank-26B-A4BMoE | I1-Q2_K | 26.5B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| G4-MeroMero-26B-A4B-it-uncensored-hereticMoE | I1-Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| EVE-26b-XENO-HATMoE | I1-Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| Gemma-4-26B-A4B-Animus-V14.1-FFT-hereticMoE | I1-Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| gemma-4-26B-A4B-it-qat-q4_0-unquantized-hereticMoE | I1-Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliteratedMoE | I1-Q2_K | 26.5B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| G4-MeroMero-26B-A4BMoE | I1-Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| G4-Dark-Soul-26B-A4BMoE | I1-Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| gemma-4-26B-A4B-it-local-abliterated-sota-internal-t34MoE | I1-Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
| gemma-4-26B-A4B-it-SOMPOA-heresyMoE | I1-Q2_K | 25.8B | 9.86 GiB | 0.49 GiB | 11.13 GiB | 0.03 GiB | 38±12.9% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Questions people ask
- What AI models can a Titan Xp run?
- 1714 of 2118 indexed open-weight models fit a Titan Xp at 16,384 context with q8_0 KV cache, the largest being DeepSeek-V2-Lite-Chat at Q5_0. That covers text, vision-language, image, video and speech models.
- How much usable memory does a Titan Xp actually have?
- Its nameplate is 12 GB, but about 11.16 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a Titan Xp fast for local AI?
- Its memory bandwidth is 548 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.