AMD · workstation

Radeon AI Pro R9600D

Radeon AI Pro R9600D has 32 GB of VRAM at 640 GB/s — about 29.76 GiB usable after driver and compositor overhead. 1859 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
32 GB
GDDR6
Bandwidth
640 GB/s
256-bit bus
Tensor FP16
dense
TDP
150 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1580vision language 176video 16audio tts 21image 1audio asr 39embedding 26

What fits at 128K context

largest quantization that fits, per model · 1859 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Gemma-4-Novelist-Eclipse-31BIQ4_NL32.7B17.53 GiB11.25 GiB29.75 GiB0.01 GiB14±26.5%
Gemma-4-31B-StyleTuneIQ4_NL32.7B17.53 GiB11.25 GiB29.75 GiB0.01 GiB14±26.5%
spoomplesmaxx-v2.1-30BI1-IQ3_S28.9B11.74 GiB17.00 GiB29.75 GiB0.01 GiB14±26.5%
Huihui-granite-4.1-30b-abliteratedI1-IQ3_S28.9B11.74 GiB17.00 GiB29.75 GiB0.01 GiB14±26.5%
granite-4.1-30b-hereticI1-IQ3_S28.9B11.74 GiB17.00 GiB29.75 GiB0.01 GiB14±26.5%
Qwen3-Coder-Next-Opus-4.6-Reasoning-DistilledMoEQ2_K27.26 GiB1.59 GiB29.74 GiB0.02 GiB58±37%
Seed-OSS-36B-Instruct-biprojected-norm-preserving-abliteratedI1-Q2_K_S36.2B11.74 GiB17.00 GiB29.74 GiB0.02 GiB14±26.5%
Hermes-4.3-36B-hereticI1-Q2_K_S36.2B11.74 GiB17.00 GiB29.74 GiB0.02 GiB14±26.5%
Skyfall-31B-v4.2Q3_K_M31.4B14.37 GiB14.34 GiB29.73 GiB0.03 GiB14±26.5%
GLM-Z1-Rumination-32B-0414Q2_K_L33.1B12.53 GiB16.20 GiB29.73 GiB0.03 GiB14±26.5%
granite-4.1-30bQ3_K_S28.9B11.71 GiB17.00 GiB29.72 GiB0.04 GiB14±26.5%
Smilodon-9B-v1F1610.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
bella-bartender-v2F169.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedF169.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
Gemma-2-9B-It-SPPO-Iter3BF169.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
G2-Darkest-Writer-Dirty-Shirley-9B-v2F169.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
G2-Darkest-Writer-9B-v1F169.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
Gemma-SEA-LION-v3-9B-ITF169.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
Tiger-Gemma-9B-v3F169.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
gemma-2-9b-itF169.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
magnum-v4-9bF169.2B17.22 GiB11.55 GiB29.71 GiB0.05 GiB14±26.5%
Seed-OSS-36B-InstructIQ2_M36.2B11.68 GiB17.00 GiB29.68 GiB0.08 GiB14±26.5%
Hermes-4.3-36BIQ2_M36.2B11.68 GiB17.00 GiB29.68 GiB0.08 GiB14±26.5%
Devstral-Small-2-24B-Instruct-2512Q6_K24.0B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Transformed-Journey-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Magistry-24B-v1.1I1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Mergedonia-AETHER-24B-v1aI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Mergedonia-AETHER-24B-v1bI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Slimaki-Tavern-24B-v1.3I1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Maginum-Cydoms-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Maginum-Cydoms-24B-absolute-heresyI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Dolphin3.0-Mistral-24BQ6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Dolphin3.0-R1-Mistral-24BQ6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Cydonia_VistralQ6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Mistral-Small-3.2-24B-Instruct-2506-ultra-uncensored-hereticI1-Q6_K24.0B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Huihui-Mistral-Small-3.2-24B-Instruct-2506-abliterated-llamacppfixedI1-Q6_K24.0B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Dans-PersonalityEngine-V1.2.0-24bI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Mistral-Small-3_2-24B-Instruct-2506-antislop.v2I1-Q6_K24.0B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Mistral-Small-3.2-24B-Instruct-2506Q6_K24.0B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Dans-PersonalityEngine-V1.3.0-24bI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Devstral-Small-2507Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Goetia-24B-v1.1I1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Devstral-Small-2505Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
MS3.2-PaintedFantasy-v3-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
RP-Spectrum-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
MS3.2-PaintedFantasy-v4.1-24B-ultra-uncensored-heretic-v2I1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Magidonia-24B-v4.3-heretic-v1.2I1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Magidonia-24B-v4.3-absolute-heresyI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
MagiSeek-Pro-V1I1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Magistral-Small-2509Q6_K24.0B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Magistral-Small-2507Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Cogidonia-v2-24BI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Magidonia-24B-v4.3I1-Q6_K18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Precog-24B-v1I1-Q6_K18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
experiment024bI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Magidonia-24B-v4.2.0Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Berthier-Mistral-Military-24BI1-Q6_K24.0B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
MS-2501-DPE-QwQify-v0.1-24BQ6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Mistral-Small-3.2-24B-Instruct-2506-llamacppfixedI1-Q6_K24.0B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
Cydonia-24B-v4.3-absolute-heresyI1-Q6_K23.6B18.02 GiB10.63 GiB29.66 GiB0.10 GiB14±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Radeon AI Pro R9600D run?
1859 of 2118 indexed open-weight models fit a Radeon AI Pro R9600D at 131,072 context with q8_0 KV cache, the largest being Gemma-4-Novelist-Eclipse-31B at IQ4_NL. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon AI Pro R9600D actually have?
Its nameplate is 32 GB, but about 29.76 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon AI Pro R9600D fast for local AI?
Its memory bandwidth is 640 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.