RTX PRO 4000 Blackwell
RTX PRO 4000 Blackwell has 24 GB of VRAM at 672 GB/s — about 22.32 GiB usable after driver and compositor overhead. 1872 of 2118 indexed models fit at 32K context with f16 KV.
What fits at 32K context
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| EXAONE-4.0-32B | Q4_K_L | 32.0B | 18.38 GiB | 2.84 GiB | 22.32 GiB | 0.00 GiB | 18±22% |
| deepseek-coder-33b-instruct | IQ3_S | 33.3B | 13.49 GiB | 7.75 GiB | 22.32 GiB | 0.00 GiB | 18±22% |
| Qwen3-42B-A3B-2507-Thinking-Abliterated-uncensored-TOTAL-RECALL-v2-Medium-MASTER-CODERMoE | I1-Q3_K_S | 42.4B | 17.13 GiB | 4.19 GiB | 22.32 GiB | 0.00 GiB | 35±37% |
| Qwen3-Coder-Next-REAMMoE | Q2_K | 60.3B | 20.56 GiB | 0.75 GiB | 22.30 GiB | 0.02 GiB | 84±37% |
| Magistry-24B-v1.1 | Q5_K_L | 23.6B | 16.18 GiB | 5.00 GiB | 22.30 GiB | 0.02 GiB | 18±22% |
| magnum-v4-27b | Q4_0 | 27.2B | 14.60 GiB | 6.56 GiB | 22.30 GiB | 0.02 GiB | 18±22% |
| Yi-34B-200K-DARE-megamerge-v8 | IQ3_XS | 34.4B | 13.71 GiB | 7.50 GiB | 22.29 GiB | 0.03 GiB | 18±22% |
| Nous-Hermes-2-Yi-34B | I1-IQ3_XS | 34.4B | 13.71 GiB | 7.50 GiB | 22.29 GiB | 0.03 GiB | 18±22% |
| WizardCoder-Python-34B-V1.0 | I1-Q3_K_M | 33.7B | 15.19 GiB | 6.00 GiB | 22.28 GiB | 0.04 GiB | 18±22% |
| Phind-CodeLlama-34B-Python-v1 | I1-Q3_K_M | 33.7B | 15.19 GiB | 6.00 GiB | 22.28 GiB | 0.04 GiB | 18±22% |
| Phind-CodeLlama-34B-v2 | I1-Q3_K_M | 33.7B | 15.19 GiB | 6.00 GiB | 22.28 GiB | 0.04 GiB | 18±22% |
| Qwen3-48B-A4B-Savant-Commander-Distill-12X-Closed-Open-Heretic-UncensoredMoE | I1-IQ4_XS | 33.6B | 16.76 GiB | 4.50 GiB | 22.27 GiB | 0.05 GiB | 27±37% |
| CodeLlama-34b-instruct-hf | Q3_K_M | 33.7B | 15.17 GiB | 6.00 GiB | 22.26 GiB | 0.06 GiB | 18±22% |
| WizardLM-1.0-Uncensored-CodeLlama-34b | Q3_K_M | 33.7B | 15.17 GiB | 6.00 GiB | 22.26 GiB | 0.06 GiB | 18±22% |
| deepseek-coder-33b-base | Q3_K_S | 33.3B | 13.43 GiB | 7.75 GiB | 22.26 GiB | 0.06 GiB | 18±22% |
| WhiteRabbitNeo-33B-v1 | Q3_K_S | 33.3B | 13.43 GiB | 7.75 GiB | 22.26 GiB | 0.06 GiB | 18±22% |
| GLM-4.7-Flash-hereticMoE | Q5_K_S | 29.9B | 19.59 GiB | 1.65 GiB | 22.26 GiB | 0.06 GiB | 54±37% |
| gemma-2-27b-it | IQ4_NL | 27.2B | 14.56 GiB | 6.56 GiB | 22.25 GiB | 0.07 GiB | 18±22% |
| Seed-OSS-36B-Instruct-biprojected-norm-preserving-abliterated | I1-IQ3_XXS | 36.2B | 13.15 GiB | 8.00 GiB | 22.25 GiB | 0.07 GiB | 18±22% |
| Seed-OSS-36B-Instruct | IQ3_XXS | 36.2B | 13.15 GiB | 8.00 GiB | 22.25 GiB | 0.07 GiB | 18±22% |
| Hermes-4.3-36B-heretic | I1-IQ3_XXS | 36.2B | 13.15 GiB | 8.00 GiB | 22.25 GiB | 0.07 GiB | 18±22% |
| Hermes-4.3-36B | IQ3_XXS | 36.2B | 13.15 GiB | 8.00 GiB | 22.25 GiB | 0.07 GiB | 18±22% |
| Qwen-AgentWorld-35B-A3BMoE | UD-Q4_K_M | 34.7B | 20.61 GiB | 0.63 GiB | 22.24 GiB | 0.08 GiB | 85±37% |
| Ornith-1.0-35BMoE | UD-Q4_K_M | 34.7B | 20.61 GiB | 0.63 GiB | 22.24 GiB | 0.08 GiB | 85±37% |
| North-Mini-Code-1.0MoE | UD-Q5_K_S | 30.5B | 20.13 GiB | 1.13 GiB | 22.24 GiB | 0.08 GiB | 61±37% |
| gemma-4-A4B-98e-v6-coder-itMoE | Q8_0 | 20.5B | 19.71 GiB | 1.54 GiB | 22.24 GiB | 0.08 GiB | 18±22% |
| gemma-4-A4B-98e-v7-coder-itMoE | Q8_0 | 20.5B | 19.71 GiB | 1.54 GiB | 22.24 GiB | 0.08 GiB | 18±22% |
| gemma-4-A4B-98e-v7-coderx-itMoE | Q8_0 | 20.5B | 19.71 GiB | 1.54 GiB | 22.24 GiB | 0.08 GiB | 18±22% |
| internlm2-math-plus-20b | I1-Q6_K | 19.9B | 15.18 GiB | 6.00 GiB | 22.24 GiB | 0.08 GiB | 18±22% |
| Skyfall-31B-v4.2 | Q3_K_M | 31.4B | 14.37 GiB | 6.75 GiB | 22.24 GiB | 0.08 GiB | 18±22% |
| MathCoder2-CodeLlama-7B | Q6_K_L | 6.7B | 5.21 GiB | 16.00 GiB | 22.23 GiB | 0.09 GiB | 18±22% |
| gemma-4-26B-A4B-itMoE | UD-Q5_K_M | 26.5B | 19.70 GiB | 1.54 GiB | 22.23 GiB | 0.09 GiB | 18±22% |
| Huihui-gemma-4-26B-A4B-it-abliteratedMoE | UD-Q5_K_M | 26.5B | 19.70 GiB | 1.54 GiB | 22.23 GiB | 0.09 GiB | 18±22% |
| Nex-N2-miniMoE | UD-Q4_K_M | 35.1B | 20.60 GiB | 0.63 GiB | 22.22 GiB | 0.10 GiB | 85±37% |
| ALIA-40b-fc-2606 | I1-IQ3_XXS | 40.4B | 15.11 GiB | 6.00 GiB | 22.22 GiB | 0.10 GiB | 18±22% |
| ALIA-40b-instruct-2606 | I1-IQ3_XXS | 40.4B | 15.11 GiB | 6.00 GiB | 22.22 GiB | 0.10 GiB | 18±22% |
| umt5-xxl | F32 | 5.7B | 21.17 GiB | 0.00 GiB | 22.22 GiB | 0.10 GiB | 18±22% |
| Phi-3-mini-4k-instructKV unresolved | Q5_K_M | 3.8B | 9.20 GiB | 12.00 GiB | 22.21 GiB | 0.11 GiB | 18±22% |
| granite-4.0-h-smallMoE | Q5_0 | 32.2B | 20.72 GiB | 0.50 GiB | 22.21 GiB | 0.11 GiB | 48±37% |
| OLMo-2-0325-32B | Q3_K_S | 32.2B | 13.09 GiB | 8.00 GiB | 22.19 GiB | 0.13 GiB | 18±22% |
| Llama3.2-30B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoE | IQ4_XS | 30.0B | 15.30 GiB | 5.88 GiB | 22.18 GiB | 0.14 GiB | 25±37% |
| Ornith-Agents-A1-3.7-35B-A3B-dare_ties_v4MoE | Q4_K_M | 34.7B | 20.55 GiB | 0.63 GiB | 22.18 GiB | 0.14 GiB | 85±37% |
| Ornith-Agents-A1-3.6-35B-A3B-dare_tiesMoE | Q4_K_M | 34.7B | 20.55 GiB | 0.63 GiB | 22.18 GiB | 0.14 GiB | 85±37% |
| deepseek-coder-6.7b-instruct | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.18 GiB | 0.14 GiB | 18±22% |
| deepseek-coder-6.7b-base | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.18 GiB | 0.14 GiB | 18±22% |
| deepseek-coder-6.7B-kexer | I1-Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.18 GiB | 0.14 GiB | 18±22% |
| Magicoder-S-DS-6.7B | I1-Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.18 GiB | 0.14 GiB | 18±22% |
| CodeLlama-7b-instruct-hf | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| CodeLlama-7b-hf | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| WizardLM-7B-Uncensored | I1-Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| Llama-2-7B-32K-Instruct | I1-Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| Luna-AI-Llama2-Uncensored | I1-Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| Llama-2-7b-chat-hf | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| Swallow-7b-NVE-instruct-hf | I1-Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| llava-v1.5-7b | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| CodeLlama-7b-python-hf | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| Wizard-Vicuna-7B-Uncensored | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| llama2_7b_chat_uncensored | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| WizardLM-7B-V1.0-Uncensored | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
| Llama-2-7b-hf | Q6_K | 6.7B | 5.15 GiB | 16.00 GiB | 22.17 GiB | 0.15 GiB | 18±22% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Measured on this card
| Workload◍ | Median | Middle 50% | Runs |
|---|---|---|---|
| Prompt processing | 4056.81 tok/s | 3153.61–5016.71 | 16 |
| Text generation | 126.09 tok/s | 117.58–132.09 | 12 |
Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-15013.
Questions people ask
- What AI models can a RTX PRO 4000 Blackwell run?
- 1872 of 2118 indexed open-weight models fit a RTX PRO 4000 Blackwell at 32,768 context with f16 KV cache, the largest being EXAONE-4.0-32B at Q4_K_L. That covers text, vision-language, image, video and speech models.
- How much usable memory does a RTX PRO 4000 Blackwell actually have?
- Its nameplate is 24 GB, but about 22.32 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a RTX PRO 4000 Blackwell fast for local AI?
- Its memory bandwidth is 672 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.