Apple · apple

Apple M4 Max

Apple M4 Max has 36 GB of unified memory at 410 GB/s — about 25.11 GiB usable after driver and compositor overhead. 2008 of 2118 indexed models fit at 16K context with q4_0 KV. Note only 27 GB of its 36 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
36 GB
LPDDR5X-8533
Bandwidth
410 GB/s
384-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1725vision language 179audio tts 21audio asr 39video 16image 2embedding 26

What fits at 16K context

largest quantization that fits, per model · 2008 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Hypernova-60B-2605MoEI1-IQ2_XS58.7B26.29 GiB0.15 GiB26.98 GiB0.02 GiB48±37%
magnum-v2-32bQ6_K_L32.5B25.20 GiB1.13 GiB26.97 GiB0.03 GiB12±8.3%
Huihui-Qwen3-Coder-Next-abliteratedMoEQ2_K_L79.7B26.29 GiB0.11 GiB26.93 GiB0.07 GiB59±37%
Qwen3-53B-A3B-2507-THINKING-TOTAL-RECALL-v2-MASTER-CODERMoEI1-Q3_K_L53.0B25.64 GiB0.74 GiB26.92 GiB0.08 GiB37±37%
Qwen3-48B-A4B-Savant-Commander-Distill-12X-Closed-Open-Heretic-UncensoredMoEI1-Q6_K33.6B25.69 GiB0.63 GiB26.88 GiB0.12 GiB27±37%
Gemma4-Gutenberg-31BQ6_K_L31.3B25.21 GiB1.03 GiB26.87 GiB0.13 GiB12±8.3%
gemma-4-31B-itQ6_K_L31.3B25.21 GiB1.03 GiB26.87 GiB0.13 GiB12±8.3%
Gemma4-Gutenberg-31B-HereticQ6_K_L31.3B25.21 GiB1.03 GiB26.87 GiB0.13 GiB12±8.3%
Equinox-31BQ6_K_L31.3B25.21 GiB1.03 GiB26.87 GiB0.13 GiB12±8.3%
gemma-4-31B-it-SDFT-Heretic-RPQ6_K_L30.7B25.21 GiB1.03 GiB26.87 GiB0.13 GiB12±8.3%
Gemma-The-Writer-N-Restless-Quill-10B-UncensoredQ6_K10.0B25.24 GiB1.04 GiB26.87 GiB0.13 GiB12±8.3%
L3-Dark-Planet-8BQ8_08.0B25.70 GiB0.56 GiB26.85 GiB0.15 GiB12±8.3%
Hunyuan-A13B-InstructMoEUD-IQ2_XXS80.4B25.73 GiB0.56 GiB26.84 GiB0.16 GiB12±8.3%
CodeLlama-70b-Instruct-hfI1-IQ3_XXS69.0B24.76 GiB1.41 GiB26.84 GiB0.16 GiB12±8.3%
CodeLlama-70b-Python-hfI1-IQ3_XXS69.0B24.76 GiB1.41 GiB26.84 GiB0.16 GiB12±8.3%
Nous-Hermes-Llama2-70bI1-IQ3_XXS69.0B24.76 GiB1.41 GiB26.84 GiB0.16 GiB12±8.3%
Midnight-Miqu-70B-v1.5I1-IQ3_XXS69.0B24.76 GiB1.41 GiB26.83 GiB0.17 GiB12±8.3%
CallerQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Dumpling-Qwen2.5-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
OREAL-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Baichuan-M2-32B-abliteratedQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
QwQ-32B-Preview-abliterated-linear25I1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
openhands-lm-32b-v0.1I1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Qwen2.5-Coder-32B-abliteratedI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
INTELLECT-2Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
LongWriter-Zero-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
m1-32bI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
XMainframe-v2-Instruct-32bI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Qwen2.5-Coder-32B-Python-SpecialistI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Qwen2.5-32b-RP-InkI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
OpenCodeReasoning-Nemotron-32B-IOIQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Qwen2.5-Coder-32B-Instruct-abliteratedQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
OlympicCoder-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
OpenCodeReasoning-Nemotron-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
OpenThinker-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
QwQ-32B-ArliAI-RpR-v4Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Qwen2.5-Coder-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
QwQ-32B-abliteratedQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
DeepSeek-R1-Distill-Qwen-32B-hereticI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
InnoSpark-HPC-RM-32BI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
OpenThinker2-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Qwen2.5-32B-InstructQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Qwen2.5-Coder-32B-Instruct-UncensoredI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
QwQ-32B-PreviewQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
DeepSeek-R1-Distill-Qwen-32B-abliteratedQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
TinyR1-32B-PreviewQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
deepseek-r1-qwen-2.5-32B-ablatedQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Rombos-LLM-V2.5-Qwen-32bQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
QwQ-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Qwen2.5-32B-ArliAI-RPMax-v1.3Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
DeepSeek-R1-Distill-Qwen-32B-Blunt-UncensoredQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
DeepSeek-R1-Distill-Qwen-32BQ6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
Qwen2.5-VL-32B-InstructQ6_K33.5B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
EVA-Qwen2.5-32B-v0.2Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
EVA-Qwen2.5-32B-v0.1Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
cogito-v1-preview-qwen-32BI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
QwQ-32B-Snowdrop-v0I1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
DeepSeek-R1-Distill-Qwen-32B-UncensoredI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
RoguePlanet-DeepSeek-R1-Qwen-32B-RPI1-Q6_K32.8B25.04 GiB1.13 GiB26.81 GiB0.19 GiB12±8.3%
OpenBuddy-R1-0528-Distill-Qwen3-32B-Preview0-QATQ6_K32.8B25.04 GiB1.13 GiB26.80 GiB0.20 GiB12±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M4 Max run?
2008 of 2118 indexed open-weight models fit a Apple M4 Max at 16,384 context with q4_0 KV cache, the largest being Hypernova-60B-2605 at I1-IQ2_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M4 Max actually have?
Its nameplate is 36 GB, but about 25.11 GiB is available to a model once driver and compositor overhead is accounted for, and only 27 GB of the pool can be allocated to the GPU at all.
Is a Apple M4 Max fast for local AI?
Its memory bandwidth is 410 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.