AMD · unified x86

AMD Ryzen AI Max 385 (Radeon 8050S)

AMD Ryzen AI Max 385 (Radeon 8050S) has 32 GB of VRAM at 256 GB/s — about 22.32 GiB usable after driver and compositor overhead. 1752 of 2118 indexed models fit at 128K context with q8_0 KV. Note only 24 GB of its 32 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
32 GB
LPDDR5X-8000
Bandwidth
256 GB/s
256-bit bus
Tensor FP16
dense
TDP
120 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1481audio asr 39vision language 168audio tts 21image 1video 16embedding 26

What fits at 128K context

largest quantization that fits, per model · 1752 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Olmo-3.1-32B-InstructQ4_K_L32.2B18.50 GiB4.70 GiB23.99 GiB0.01 GiB6±25%
Olmo-3.1-32B-ThinkQ4_K_L32.2B18.50 GiB4.70 GiB23.99 GiB0.01 GiB6±25%
Olmo-3-32B-ThinkQ4_K_L32.2B18.50 GiB4.70 GiB23.99 GiB0.01 GiB6±25%
Skyfall-31B-v4.2-hereticI1-IQ2_XS31.4B8.83 GiB14.34 GiB23.99 GiB0.01 GiB6±25%
Skyfall-31B-v4.2I1-IQ2_XS31.4B8.83 GiB14.34 GiB23.99 GiB0.01 GiB6±25%
Devstral-Small-2-24B-Instruct-2512IQ4_NL24.0B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Voxtral-Small-24B-2507IQ4_NL24.3B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Morax-24B-v2IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Dolphin3.0-R1-Mistral-24BIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Dolphin3.0-Mistral-24BIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Dans-PersonalityEngine-V1.2.0-24bIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Cydonia_VistralIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Mistral-Small-3.2-24B-Instruct-2506IQ4_NL24.0B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Dans-PersonalityEngine-V1.3.0-24bIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Devstral-Small-2507IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Devstral-Small-2505IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
MS3.2-PaintedFantasy-v3-24BIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Magistral-Small-2509IQ4_NL24.0B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Magistral-Small-2507IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Precog-24B-v1IQ4_NL12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Magidonia-24B-v4.3IQ4_NL12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Magidonia-24B-v4.2.0IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
MS-2501-DPE-QwQify-v0.1-24BIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
sarvam-mIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Magistral-Small-2506IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Cydonia-24B-v4.1IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Cydonia-24B-v4IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Mistral-Small-3.1-24B-Instruct-2503IQ4_NL24.0B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Cydonia-24B-v4.3IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Cydonia-24B-v4.2.0IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
MS3.2-24B-Magnum-DiamondIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Mistral-Small-24B-Instruct-2501-abliteratedIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Dolphin-Mistral-24B-Venice-EditionIQ4_NL24.0B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Mistral-Small-24B-Instruct-2501IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Hearthfire-24BIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Mistral-Small-24B-ArliAI-RPMax-v1.4IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Codex-24B-Small-3.2IQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
Harbinger-24BIQ4_NL23.6B12.54 GiB10.63 GiB23.99 GiB0.01 GiB6±25%
GLM-4-32B-0414-Korean-CultureI1-Q4_132.6B19.14 GiB4.05 GiB23.98 GiB0.02 GiB6±25%
GLM-Z1-32B-0414Q4_132.6B19.14 GiB4.05 GiB23.98 GiB0.02 GiB6±25%
GLM-4-32B-0414Q4_132.6B19.14 GiB4.05 GiB23.98 GiB0.02 GiB6±25%
Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingI1-IQ3_M39.5B16.83 GiB6.38 GiB23.97 GiB0.03 GiB6±25%
Qwen3.5-40B-RoughHouse-Claude-4.6-Opus-Polar-Deckard-Uncensored-Heretic-ThinkingI1-IQ3_M39.5B16.83 GiB6.38 GiB23.97 GiB0.03 GiB6±25%
Slimaki-Tavern-24B-v1.3Q4_023.6B12.52 GiB10.63 GiB23.96 GiB0.04 GiB6±25%
mistral-small-3.1-24b-instruct-2503-hfQ4_023.6B12.52 GiB10.63 GiB23.96 GiB0.04 GiB6±25%
Qwen3.6-34B-80L-Fable-5-HereticI1-Q4_033.4B17.88 GiB5.31 GiB23.96 GiB0.04 GiB6±25%
NVIDIA-Nemotron-Nano-12B-v2Q4_K_S12.3B6.71 GiB16.47 GiB23.95 GiB0.05 GiB6±25%
Snowpiercer-15B-v4-hereticI1-Q5_K_M15.0B9.92 GiB13.28 GiB23.95 GiB0.05 GiB6±25%
Snowpiercer-15B-v4Q5_K_M15.0B9.92 GiB13.28 GiB23.95 GiB0.05 GiB6±25%
North-Mini-Code-1.0MoEUD-Q5_K_M30.5B21.37 GiB1.89 GiB23.94 GiB0.06 GiB18±37%
Qwen3.6-27B-Omnimerge-v4Q5_K_L27.8B18.92 GiB4.25 GiB23.94 GiB0.06 GiB6±25%
Phi-4-reasoningQ5_K_M14.7B9.88 GiB13.28 GiB23.92 GiB0.08 GiB6±25%
Phi-4-reasoning-plusQ5_K_M14.7B9.88 GiB13.28 GiB23.92 GiB0.08 GiB6±25%
phi-4Q5_K_M14.7B9.88 GiB13.28 GiB23.92 GiB0.08 GiB6±25%
Nemotron-Labs-Audex-30B-A3BQ4_K_M32.0B23.16 GiB0.00 GiB23.90 GiB0.10 GiB6±25%
Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-hereticQ5_127.4B18.89 GiB4.25 GiB23.90 GiB0.10 GiB6±25%
medgemma-27b-itI1-Q5_K_S28.8B17.48 GiB5.64 GiB23.90 GiB0.10 GiB6±25%
gemma-3-27b-it-abliterated-refined-visionI1-Q5_K_S27.4B17.48 GiB5.64 GiB23.90 GiB0.10 GiB6±25%
gemma-3-27b-it-abliteratedQ5_K_S27.4B17.48 GiB5.64 GiB23.90 GiB0.10 GiB6±25%
Nidum-Gemma-3-27B-it-UncensoredI1-Q5_K_S27.4B17.48 GiB5.64 GiB23.90 GiB0.10 GiB6±25%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a AMD Ryzen AI Max 385 (Radeon 8050S) run?
1752 of 2118 indexed open-weight models fit a AMD Ryzen AI Max 385 (Radeon 8050S) at 131,072 context with q8_0 KV cache, the largest being Olmo-3.1-32B-Instruct at Q4_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a AMD Ryzen AI Max 385 (Radeon 8050S) actually have?
Its nameplate is 32 GB, but about 22.32 GiB is available to a model once driver and compositor overhead is accounted for, and only 24 GB of the pool can be allocated to the GPU at all.
Is a AMD Ryzen AI Max 385 (Radeon 8050S) fast for local AI?
Its memory bandwidth is 256 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.