NVIDIA · workstation

RTX PRO 5000 Blackwell

RTX PRO 5000 Blackwell has 72 GB of VRAM at 1344 GB/s — about 66.96 GiB usable after driver and compositor overhead. 2070 of 2118 indexed models fit at 64K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
72 GB
GDDR7
Bandwidth
1344 GB/s
384-bit bus
Tensor FP16
295 TF
dense
TDP
300 W
$4569 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1780vision language 186image 2audio asr 39audio tts 21video 16embedding 26

What fits at 64K context

largest quantization that fits, per model · 2070 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEIQ4_NL124B64.39 GiB1.55 GiB66.93 GiB0.03 GiB53±37%
GLM-4.6VMoEQ4_1108B62.66 GiB3.23 GiB66.92 GiB0.04 GiB40±37%
command-a-plus-05-2026-bf16MoEIQ2_S219B65.10 GiB0.68 GiB66.79 GiB0.17 GiB51±37%
Ornith-Agents-A1-3.7-35B-A3B-dare_ties_v4MoEF1634.7B65.34 GiB0.35 GiB66.70 GiB0.26 GiB69±37%
Ornith-Agents-A1-3.6-35B-A3B-dare_tiesMoEF1634.7B65.34 GiB0.35 GiB66.70 GiB0.26 GiB69±37%
HuatuoGPT-o1-72BQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Rombo-LLM-V3.0-Qwen-72bQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Qwen2.5-72B-Instruct-abliteratedQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
EVA-Qwen2.5-72B-v0.2Q6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
MiroThinker-v1.0-72BQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Qwen2.5-Math-72B-InstructQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Qwen2.5-72B-InstructQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Qwen2.5-72BQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Kimi-Dev-72BQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
magnum-v4-72bQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
KAT-Dev-72B-ExpQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Chuluun-Qwen2.5-72B-v0.01Q6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Homer-v1.0-Qwen2.5-72BQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Qwen2.5-VL-72B-InstructQ6_K73.4B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Tower-Plus-72B-ultra-uncensored-hereticI1-Q6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
Chronos-Platinum-72BQ6_K72.7B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
UI-TARS-72B-DPOQ6_K73.4B59.93 GiB5.63 GiB66.68 GiB0.28 GiB12±22%
GLM-4.5VMoEI1-Q4_1108B62.40 GiB3.23 GiB66.66 GiB0.30 GiB41±37%
CallerBF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
Dumpling-Qwen2.5-32BBF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
OpenThinker-32BF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
INTELLECT-2BF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
openhands-lm-32b-v0.1BF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
LongWriter-Zero-32BBF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
OpenCodeReasoning-Nemotron-32BBF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
OpenCodeReasoning-Nemotron-32B-IOIBF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
Qwen2.5-Coder-32B-Instruct-abliteratedF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
QwQ-32B-ArliAI-RpR-v4BF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
OpenThinker2-32BBF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
Qwen2.5-Coder-32BF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
Qwen2.5-32B-InstructF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
QwQ-32B-PreviewBF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
Qwen2.5-32b-RP-InkF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
deepseek-r1-qwen-2.5-32B-ablatedBF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
Rombos-LLM-V2.5-Qwen-32bF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
DeepSeek-R1-Distill-Qwen-32B-abliteratedBF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
Qwen2.5-32B-ArliAI-RPMax-v1.3F1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
DeepSeek-R1-Distill-Qwen-32BF1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
Qwen2.5-VL-32B-InstructBF1633.5B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
EVA-Qwen2.5-32B-v0.2F1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
EVA-Qwen2.5-32B-v0.1F1632.8B61.04 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
cogito-v1-preview-qwen-32BBF1632.8B61.03 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
QwQ-32B-Snowdrop-v0BF1632.8B61.03 GiB4.50 GiB66.63 GiB0.33 GiB12±22%
OpenBuddy-R1-0528-Distill-Qwen3-32B-Preview0-QATBF1632.8B61.03 GiB4.50 GiB66.62 GiB0.34 GiB12±22%
KAT-DevBF1632.8B61.03 GiB4.50 GiB66.62 GiB0.34 GiB12±22%
Qwen3-VL-32B-InstructBF1633.4B61.03 GiB4.50 GiB66.62 GiB0.34 GiB12±22%
Qwen3-VL-32B-ThinkingBF1633.4B61.03 GiB4.50 GiB66.62 GiB0.34 GiB12±22%
Qwen3-VL-32B-Instruct-ultra-uncensored-hereticBF1633.4B61.03 GiB4.50 GiB66.62 GiB0.34 GiB12±22%
Qwen3-32BBF1632.8B61.03 GiB4.50 GiB66.62 GiB0.34 GiB12±22%
DeepSWE-PreviewBF1632.8B61.03 GiB4.50 GiB66.62 GiB0.34 GiB12±22%
Llama-3_3-Nemotron-Super-49B-v1_5Q3_K_S49.9B20.45 GiB45.00 GiB66.59 GiB0.37 GiB12±22%
Valkyrie-49B-v2.1I1-IQ3_S49.9B20.45 GiB45.00 GiB66.59 GiB0.37 GiB12±22%
Llama-3_3-Nemotron-Super-49B-v1Q3_K_S49.9B20.45 GiB45.00 GiB66.59 GiB0.37 GiB12±22%
Step-3.5-Flash-REAP-121B-A11BI1-Q3_K_L121B58.48 GiB7.04 GiB66.55 GiB0.41 GiB12±22%
Mistral-Small-4-119B-2603MoEQ4_K_S119B65.08 GiB0.40 GiB66.51 GiB0.45 GiB68±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX PRO 5000 Blackwell run?
2070 of 2118 indexed open-weight models fit a RTX PRO 5000 Blackwell at 65,536 context with q4_0 KV cache, the largest being NVIDIA-Nemotron-3-Super-120B-A12B-BF16 at IQ4_NL. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX PRO 5000 Blackwell actually have?
Its nameplate is 72 GB, but about 66.96 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX PRO 5000 Blackwell fast for local AI?
Its memory bandwidth is 1344 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.