AMD · unified x86

AMD Ryzen AI Max+ 395 (Radeon 8060S)

AMD Ryzen AI Max+ 395 (Radeon 8060S) has 32 GB of VRAM at 256 GB/s — about 22.32 GiB usable after driver and compositor overhead. 1881 of 2118 indexed models fit at 128K context with q4_0 KV. Note only 24 GB of its 32 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
32 GB
LPDDR5X-8000
Bandwidth
256 GB/s
256-bit bus
Tensor FP16
dense
TDP
120 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1604vision language 173image 2audio tts 21audio asr 39video 16embedding 26

What fits at 128K context

largest quantization that fits, per model · 1881 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
c4ai-command-r-08-2024Q4_K_S32.3B17.55 GiB5.63 GiB23.99 GiB0.01 GiB6±25%
Olmo-3.1-32B-InstructQ5_K_S32.2B20.71 GiB2.49 GiB23.99 GiB0.01 GiB6±25%
Olmo-3.1-32B-ThinkQ5_K_S32.2B20.71 GiB2.49 GiB23.99 GiB0.01 GiB6±25%
Olmo-3-32B-ThinkQ5_K_S32.2B20.71 GiB2.49 GiB23.99 GiB0.01 GiB6±25%
MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-ThinkingI1-IQ4_XS23.4B11.84 GiB11.39 GiB23.98 GiB0.02 GiB6±25%
Hy-MT2-30B-A3BMoEQ5_K_M30.1B19.91 GiB3.38 GiB23.98 GiB0.02 GiB14±37%
Gemma-4-31B-Isometry-RPI1-Q4_032.7B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
Gemma-4-Dark-Gemistry-31BI1-Q4_032.7B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
Prosopon-31BI1-Q4_032.7B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
Gemma-4-Novelist-Eclipse-31BI1-Q4_032.7B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
Giftige-Blume-31B-v1-StyleSwapI1-Q4_032.7B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
G4-MeroMero-31B-StyleSwapI1-Q4_032.7B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
Gemma-4-31B-StyleTune-heretic-araI1-Q4_032.7B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
Pantheon-Reasoning-31B-1.1I1-Q4_032.7B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
Gemma-4-31B-StyleTuneI1-Q4_032.7B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
Barcenas-StyleTune-31B-FableI1-Q4_032.1B17.22 GiB5.95 GiB23.96 GiB0.04 GiB6±25%
Fallen-Gemma3-27B-v1Q6_K_L27.4B20.96 GiB2.23 GiB23.93 GiB0.07 GiB6±25%
MathCoder2-CodeLlama-7BQ6_K_L6.7B5.21 GiB18.00 GiB23.93 GiB0.07 GiB6±25%
GLM-4-32B-0414-Korean-CultureI1-Q5_K_S32.6B20.98 GiB2.14 GiB23.91 GiB0.09 GiB6±25%
GLM-Z1-32B-0414Q5_K_S32.6B20.98 GiB2.14 GiB23.91 GiB0.09 GiB6±25%
GLM-4-32B-0414Q5_K_S32.6B20.98 GiB2.14 GiB23.91 GiB0.09 GiB6±25%
GLM-Z1-32B-0414-uncensored-heretic-v2Q5_K_S32.6B20.98 GiB2.14 GiB23.91 GiB0.09 GiB6±25%
Darwin-35B-A3B-OpusMoEQ5_K_S36.0B22.50 GiB0.70 GiB23.90 GiB0.10 GiB28±37%
Aurora-Code-1MoEQ5_K_S34.7B22.50 GiB0.70 GiB23.90 GiB0.10 GiB28±37%
grug-35b-v2MoEQ5_K_S35.1B22.50 GiB0.70 GiB23.90 GiB0.10 GiB28±37%
grug-35bMoEQ5_K_S35.1B22.50 GiB0.70 GiB23.90 GiB0.10 GiB28±37%
WorldSim-Opus-3.6-35B-A3BMoEQ5_K_S35.1B22.50 GiB0.70 GiB23.90 GiB0.10 GiB28±37%
Qwen3.6-35B-A3B-AnkoMoEQ5_K_S35.1B22.50 GiB0.70 GiB23.90 GiB0.10 GiB28±37%
KAT-Coder-V2.5-DevMoEQ5_K_S34.7B22.50 GiB0.70 GiB23.90 GiB0.10 GiB28±37%
Ornith-1.0-35BMoEQ5_K_S34.7B22.50 GiB0.70 GiB23.90 GiB0.10 GiB28±37%
Nex-N2-miniMoEQ5_K_S35.1B22.50 GiB0.70 GiB23.90 GiB0.10 GiB28±37%
Nemotron-Labs-Audex-30B-A3BQ4_K_M32.0B23.16 GiB0.00 GiB23.90 GiB0.10 GiB6±25%
Pantheon-Reasoning-27BI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-PreservedI1-Q6_K27.4B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTPI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwen3.6-27B-Fable-5-ExperimentalI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwable-5-27B-CoderI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16Q6_K27.4B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
EVE-27b-XENO-HAT-DeepSeek-V4-FlashI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
EVE-27B-XENO-HATI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Godoter-27BI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Reasoning-Medical-27BI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwopus3.6-27B-v2-abliteratedI1-Q6_K27.4B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-BF16I1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Reasoning-Medical0.1-27BI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Huihui-ThinkingCap-Qwen3.6-27B-abliteratedI1-Q6_K27.4B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Semancer-27BI1-Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwen3.6-27B-Uncensored-CyberQ6_K27.4B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwen3.6-27B-Omnimerge-v4Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwopus3.6-27B-v2Q6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Darwin-28B-CoderI1-Q6_K26.9B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
Qwopus3.6-27B-CoderQ6_K27.8B20.89 GiB2.25 GiB23.90 GiB0.10 GiB6±25%
spoomplesmaxx-v2.1-30BI1-Q3_K_L28.9B14.09 GiB9.00 GiB23.90 GiB0.10 GiB6±25%
Huihui-granite-4.1-30b-abliteratedI1-Q3_K_L28.9B14.09 GiB9.00 GiB23.90 GiB0.10 GiB6±25%
granite-4.1-30b-hereticI1-Q3_K_L28.9B14.09 GiB9.00 GiB23.90 GiB0.10 GiB6±25%
granite-4.1-30bQ3_K_L28.9B14.09 GiB9.00 GiB23.90 GiB0.10 GiB6±25%
gemma-4-31B-itQ4_031.3B17.16 GiB5.95 GiB23.90 GiB0.10 GiB6±25%
deepseek-coder-6.7b-instructQ6_K6.7B5.15 GiB18.00 GiB23.88 GiB0.12 GiB6±25%
deepseek-coder-6.7b-baseQ6_K6.7B5.15 GiB18.00 GiB23.88 GiB0.12 GiB6±25%
deepseek-coder-6.7B-kexerI1-Q6_K6.7B5.15 GiB18.00 GiB23.88 GiB0.12 GiB6±25%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a AMD Ryzen AI Max+ 395 (Radeon 8060S) run?
1881 of 2118 indexed open-weight models fit a AMD Ryzen AI Max+ 395 (Radeon 8060S) at 131,072 context with q4_0 KV cache, the largest being c4ai-command-r-08-2024 at Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a AMD Ryzen AI Max+ 395 (Radeon 8060S) actually have?
Its nameplate is 32 GB, but about 22.32 GiB is available to a model once driver and compositor overhead is accounted for, and only 24 GB of the pool can be allocated to the GPU at all.
Is a AMD Ryzen AI Max+ 395 (Radeon 8060S) fast for local AI?
Its memory bandwidth is 256 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.