NVIDIA · workstation

RTX PRO 4500 Blackwell

RTX PRO 4500 Blackwell has 32 GB of VRAM at 896 GB/s — about 29.76 GiB usable after driver and compositor overhead. 2019 of 2118 indexed models fit at 16K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
32 GB
GDDR7
Bandwidth
896 GB/s
256-bit bus
Tensor FP16
dense
TDP
200 W
$2623 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
video 16vision language 181text 1734image 2embedding 26audio tts 21audio asr 39

What fits at 16K context

largest quantization that fits, per model · 2019 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Bernini-RQ8_014.3B28.71 GiB0.00 GiB29.76 GiB0.00 GiB18±22%
Salience-1.5-ProMoEQ6_K_L36.0B28.66 GiB0.09 GiB29.75 GiB0.01 GiB104±37%
Qwable-v1MoEQ6_K_L36.0B28.66 GiB0.09 GiB29.75 GiB0.01 GiB104±37%
T-SearchMoEQ6_K_L36.0B28.66 GiB0.09 GiB29.75 GiB0.01 GiB104±37%
Melody1437-27BQ3_K_M27.8B28.40 GiB0.28 GiB29.75 GiB0.01 GiB18±22%
Qwen3-53B-A3B-2507-THINKING-TOTAL-RECALL-v2-MASTER-CODERMoEI1-Q4_053.0B28.02 GiB0.74 GiB29.75 GiB0.01 GiB63±37%
Gemma-3-27B-MeditronFOQ8_028.8B28.13 GiB0.52 GiB29.73 GiB0.03 GiB18±22%
Qwen2.5-7B-Instruct-1MF327.6B28.38 GiB0.25 GiB29.68 GiB0.08 GiB18±22%
DeepSeek-R1-Distill-Qwen-7BF327.6B28.38 GiB0.25 GiB29.68 GiB0.08 GiB18±22%
UI-TARS-7B-DPOF328.3B28.38 GiB0.25 GiB29.68 GiB0.08 GiB18±22%
Qwen2-7B-InstructF327.6B28.38 GiB0.25 GiB29.68 GiB0.08 GiB18±22%
Hercules-5.0-Qwen2-7BF327.6B28.38 GiB0.25 GiB29.68 GiB0.08 GiB18±22%
Kepler-8B-Instruct-v2F167.6B28.37 GiB0.25 GiB29.67 GiB0.09 GiB18±22%
MiniCPM-o-2_6F328.7B28.37 GiB0.25 GiB29.67 GiB0.09 GiB18±22%
Mixtral-8x22B-Instruct-v0.1MoEIQ1_S141B27.61 GiB0.98 GiB29.66 GiB0.10 GiB32±37%
Mixtral-8x22B-v0.1MoEIQ1_S141B27.61 GiB0.98 GiB29.65 GiB0.11 GiB32±37%
CalmeRys-78B-Orpo-v0.1I1-IQ2_XS78.0B27.00 GiB1.51 GiB29.64 GiB0.12 GiB18±22%
calme-2.3-rys-78bIQ2_XS78.0B27.00 GiB1.51 GiB29.64 GiB0.12 GiB18±22%
Qwen3-42B-A3B-2507-Thinking-Abliterated-uncensored-TOTAL-RECALL-v2-Medium-MASTER-CODERMoEI1-Q5_K_M42.4B28.05 GiB0.59 GiB29.64 GiB0.12 GiB72±37%
Huihui-GLM-4.7-Flash-abliterated-57BMoEIQ4_XS57.3B27.93 GiB0.59 GiB29.56 GiB0.20 GiB72±37%
Assistant_Pepe_70BIQ3_XXS70.6B27.03 GiB1.41 GiB29.56 GiB0.20 GiB18±22%
Qwen3.5-88BMoEI1-Q2_K_S87.7B28.30 GiB0.11 GiB29.43 GiB0.33 GiB93±37%
DeepCoder-14B-PreviewBF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
SuperNova-MediusF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
Qwen2.5-14B-Instruct-1MF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
OpenCodeReasoning-Nemotron-14BBF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
Qwen2.5-Coder-14B-Instruct-abliteratedF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
0x-liteF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
Qwen2.5-Coder-14B-InstructF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
Qwen2.5-14B-Instruct-1M-abliteratedF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
AceReason-Nemotron-14BBF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
Qwen2.5-14B-InstructF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
Qwen2.5-Coder-14BF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
DeepSeek-R1-Distill-Qwen-14B-abliterated-v2F1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
DeepSeek-R1-Distill-Qwen-14BF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
Sugoi-14B-Ultra-HFF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
UwU-14B-Math-v0.2F1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
oxy-1-smallF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
EVA-Qwen2.5-14B-v0.2F1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
EVA-Qwen2.5-14B-v0.0F1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
EVA-Qwen2.5-14B-v0.1F1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
Strand-Rust-Coder-14B-v1BF1614.8B27.52 GiB0.84 GiB29.41 GiB0.35 GiB18±22%
Lamarck-14B-v0.7F1614.8B27.51 GiB0.84 GiB29.40 GiB0.36 GiB18±22%
command-r-35b-writer-v2I1-Q5_K_S35.0B22.67 GiB5.63 GiB29.40 GiB0.36 GiB18±22%
Kimi-Linear-48B-A3B-InstructMoEQ4_K_L49.1B28.26 GiB0.13 GiB29.40 GiB0.36 GiB18±22%
WizardLM-Uncensored-SuperCOT-StoryTelling-30bQ5_K_M32.5B21.46 GiB6.86 GiB29.39 GiB0.37 GiB18±22%
Wizard-Vicuna-30B-UncensoredI1-Q5_K_M32.5B21.46 GiB6.86 GiB29.39 GiB0.37 GiB18±22%
archangel_sft-kto_llama30bI1-Q5_K_M32.5B21.46 GiB6.86 GiB29.39 GiB0.37 GiB18±22%
Phi-3.5-MoE-instructMoEKV unresolvedQ5_K_L41.9B27.75 GiB0.56 GiB29.32 GiB0.44 GiB52±37%
Darwin-35B-A3B-OpusMoEQ6_K_L36.0B28.22 GiB0.09 GiB29.31 GiB0.45 GiB106±37%
Aurora-Code-1MoEQ6_K_L34.7B28.22 GiB0.09 GiB29.31 GiB0.45 GiB106±37%
grug-35b-v2MoEQ6_K_L35.1B28.22 GiB0.09 GiB29.31 GiB0.45 GiB106±37%
grug-35bMoEQ6_K_L35.1B28.22 GiB0.09 GiB29.31 GiB0.45 GiB106±37%
WorldSim-Opus-3.6-35B-A3BMoEQ6_K_L35.1B28.22 GiB0.09 GiB29.31 GiB0.45 GiB106±37%
Qwen3.6-35B-A3B-AnkoMoEQ6_K_L35.1B28.22 GiB0.09 GiB29.31 GiB0.45 GiB106±37%
KAT-Coder-V2.5-DevMoEQ6_K_L34.7B28.22 GiB0.09 GiB29.31 GiB0.45 GiB106±37%
Ornith-1.0-35BMoEQ6_K_L34.7B28.22 GiB0.09 GiB29.31 GiB0.45 GiB106±37%
Nex-N2-miniMoEQ6_K_L35.1B28.22 GiB0.09 GiB29.31 GiB0.45 GiB106±37%
deepseek-llm-67b-chatQ2_K67.4B26.54 GiB1.67 GiB29.31 GiB0.45 GiB18±22%
deepseek-llm-67b-baseQ2_K67.4B26.54 GiB1.67 GiB29.31 GiB0.45 GiB18±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing4623.70 tok/s3136.076649.9324
Text generation155.65 tok/s81.58169.6520
Benchmarked· n=24

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-15013.

Questions people ask

What AI models can a RTX PRO 4500 Blackwell run?
2019 of 2118 indexed open-weight models fit a RTX PRO 4500 Blackwell at 16,384 context with q4_0 KV cache, the largest being Bernini-R at Q8_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX PRO 4500 Blackwell actually have?
Its nameplate is 32 GB, but about 29.76 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX PRO 4500 Blackwell fast for local AI?
Its memory bandwidth is 896 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.