NVIDIA · workstation

RTX PRO 5000 Blackwell

RTX PRO 5000 Blackwell has 72 GB of VRAM at 1344 GB/s — about 66.96 GiB usable after driver and compositor overhead. 2063 of 2118 indexed models fit at 64K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
72 GB
GDDR7
Bandwidth
1344 GB/s
384-bit bus
Tensor FP16
295 TF
dense
TDP
300 W
$4569 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1773vision language 186image 2audio asr 39audio tts 21video 16embedding 26

What fits at 64K context

largest quantization that fits, per model · 2063 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
command-r-35b-writer-v2I1-Q5_K_M35.0B23.29 GiB42.50 GiB66.90 GiB0.06 GiB12±22%
Mistral-Small-4-119B-2603MoEQ4_K_S119B65.08 GiB0.75 GiB66.86 GiB0.10 GiB65±37%
CalmeRys-78B-Orpo-v0.1Q5_K_M78.0B54.30 GiB11.42 GiB66.85 GiB0.11 GiB12±22%
calme-2.3-rys-78bQ5_K_M78.0B54.30 GiB11.42 GiB66.85 GiB0.11 GiB12±22%
Gemma-4-Novelist-Eclipse-31BBF1632.7B59.82 GiB5.94 GiB66.84 GiB0.12 GiB12±22%
Gemma-4-31B-StyleTuneBF1632.7B59.82 GiB5.94 GiB66.84 GiB0.12 GiB12±22%
Qwen3.5-REAP-212B-A17BMoEIQ2_M212B64.76 GiB1.00 GiB66.81 GiB0.15 GiB60±37%
grok-2MoEIQ1_M270B57.16 GiB8.50 GiB66.81 GiB0.15 GiB17±37%
Wizard-Vicuna-30B-UncensoredI1-IQ3_M32.5B13.86 GiB51.80 GiB66.73 GiB0.23 GiB12±22%
archangel_sft-kto_llama30bI1-IQ3_M32.5B13.86 GiB51.80 GiB66.73 GiB0.23 GiB12±22%
Skyfall-31B-v4.2BF1631.4B58.41 GiB7.17 GiB66.70 GiB0.26 GiB12±22%
Qwen3.5-122B-A10B-hereticMoEI1-Q4_K_S123B64.86 GiB0.80 GiB66.68 GiB0.28 GiB64±37%
c4ai-command-r-08-2024F1632.3B60.17 GiB5.31 GiB66.59 GiB0.37 GiB12±22%
Laguna-S-2.1MoEUD-Q4_K_S118B63.88 GiB1.67 GiB66.57 GiB0.39 GiB55±37%
Qwen3.5-REAP-262B-A17BMoEIQ2_XXS262B64.50 GiB1.00 GiB66.55 GiB0.41 GiB64±37%
GLM-4.5-Air-DerestrictedMoEQ4_0110B59.38 GiB6.11 GiB66.52 GiB0.44 GiB33±37%
GLM-4.5-AirMoEQ4_0110B59.38 GiB6.11 GiB66.52 GiB0.44 GiB33±37%
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEIQ4_XS124B62.59 GiB2.92 GiB66.51 GiB0.45 GiB47±37%
MiMo-V2-FlashMoEKV unresolvedIQ1_M310B61.31 GiB3.98 GiB66.34 GiB0.62 GiB45±37%
Qwen3.6-35B-A3B-uncensored-hereticMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Darwin-35B-A3B-OpusMoEBF1636.0B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Carnice-MoE-35B-A3BMoEF1636.0B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Qwen3.6-35B-A3B-hereticMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Aurora-Code-1MoEBF1634.7B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
grug-35b-v2MoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
grug-35bMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
WorldSim-Opus-3.6-35B-A3BMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Qwen3.6-35B-A3B-abliterated-MAXMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Huihui-Nex-N2-mini-abliteratedMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Qwen3.6-35B-A3B-AnkoMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
KAT-Coder-V2.5-DevMoEBF1634.7B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Ornith-1.0-35B-uncensored-hereticMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Qwen3.6-35B-A3B-abliterated-v4MoEBF1634.7B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Qwen3.5-35B-A3B-ultra-uncensored-hereticMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Nex-N2-mini-ultra-uncensored-hereticMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Agents-A1MoEF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Qwen3.6-35B-A3B-java-v1MoEBF1634.7B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Nex-N2-miniMoEBF1635.1B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
Qwen3.5-35B-A3B-BaseMoEBF1636.0B64.61 GiB0.66 GiB66.28 GiB0.68 GiB66±37%
command-a-plus-05-2026-bf16MoEIQ2_XS219B63.98 GiB1.29 GiB66.27 GiB0.69 GiB48±37%
Llama-3.3-70B-InstructQ6_K_L70.6B54.39 GiB10.63 GiB66.14 GiB0.82 GiB12±22%
Anubis-70B-v1.2Q6_K_L70.6B54.39 GiB10.63 GiB66.14 GiB0.82 GiB12±22%
Tess-R1-Limerick-Llama-3.1-70BQ6_K_L70.6B54.39 GiB10.63 GiB66.14 GiB0.82 GiB12±22%
Infinity-Instruct-7M-Gen-Llama3_1-70BQ6_K_L70.6B54.39 GiB10.63 GiB66.14 GiB0.82 GiB12±22%
Athene-70BQ6_K_L70.6B54.39 GiB10.63 GiB66.14 GiB0.82 GiB12±22%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedQ4_0109B58.72 GiB6.38 GiB66.13 GiB0.83 GiB32±37%
MiniMax-M2.1-REAP-139B-A10BMoEI1-IQ3_M139B56.81 GiB8.23 GiB66.03 GiB0.93 GiB30±37%
m51Lab-MiniMax-M2.7-REAP-139B-A10BMoEI1-IQ3_M139B56.81 GiB8.23 GiB66.03 GiB0.93 GiB30±37%
step-3.5-flashIQ2_S199B51.70 GiB13.30 GiB66.03 GiB0.93 GiB12±22%
WizardLM-Uncensored-SuperCOT-StoryTelling-30bQ3_K_S32.5B13.10 GiB51.80 GiB65.97 GiB0.99 GiB12±22%
Qwen3-VL-235B-A22B-ThinkingMoEUD-IQ1_S236B58.65 GiB6.24 GiB65.92 GiB1.04 GiB33±37%
Mistral-Medium-3.5-128BQ3_K_S128B53.05 GiB11.69 GiB65.89 GiB1.07 GiB12±22%
MiniMax-M2.7MoEIQ2_XXS229B56.67 GiB8.23 GiB65.89 GiB1.07 GiB33±37%
Mixtral-8x22B-Instruct-v0.1MoEQ3_K_S141B57.28 GiB7.44 GiB65.78 GiB1.18 GiB18±37%
Mixtral-8x22B-v0.1MoEQ3_K_S141B57.28 GiB7.44 GiB65.78 GiB1.18 GiB18±37%
Mixtral-8x22B-v0.1MoEQ3_K_S141B57.27 GiB7.44 GiB65.77 GiB1.19 GiB18±37%
Qwen3-VL-235B-A22B-InstructMoEUD-IQ1_S236B58.49 GiB6.24 GiB65.77 GiB1.19 GiB33±37%
Phi-3-mini-4k-instructKV unresolvedF323.8B52.01 GiB12.75 GiB65.76 GiB1.20 GiB12±22%
Apertus-70B-Instruct-2509Q6_K70.6B53.95 GiB10.63 GiB65.75 GiB1.21 GiB12±22%
Devstral-2-123B-Instruct-2512IQ3_M125B52.89 GiB11.69 GiB65.74 GiB1.22 GiB12±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX PRO 5000 Blackwell run?
2063 of 2118 indexed open-weight models fit a RTX PRO 5000 Blackwell at 65,536 context with q8_0 KV cache, the largest being command-r-35b-writer-v2 at I1-Q5_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX PRO 5000 Blackwell actually have?
Its nameplate is 72 GB, but about 66.96 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX PRO 5000 Blackwell fast for local AI?
Its memory bandwidth is 1344 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.