Apple · apple

Apple M4 Max

Apple M4 Max has 48 GB of unified memory at 546 GB/s — about 33.48 GiB usable after driver and compositor overhead. 1996 of 2118 indexed models fit at 128K context with q4_0 KV. Note only 36 GB of its 48 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
48 GB
LPDDR5X-8533
Bandwidth
546 GB/s
512-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1712vision language 180video 16audio tts 21audio asr 39image 2embedding 26

What fits at 128K context

largest quantization that fits, per model · 1996 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Kimi-Dev-72BUD-IQ2_XXS72.7B23.94 GiB11.25 GiB35.87 GiB0.13 GiB12±8.3%
Qwen2.5-VL-72B-InstructUD-IQ2_XXS73.4B23.94 GiB11.25 GiB35.87 GiB0.13 GiB12±8.3%
EXAONE-4.5-33BQ8_034.4B32.73 GiB2.49 GiB35.86 GiB0.14 GiB12±8.3%
Yi-34B-200K-DARE-megamerge-v8Q6_K34.4B26.78 GiB8.44 GiB35.85 GiB0.15 GiB12±8.3%
Nous-Hermes-2-Yi-34BI1-Q6_K34.4B26.78 GiB8.44 GiB35.85 GiB0.15 GiB12±8.3%
Nous-Capybara-limarpv3-34BI1-Q6_K34.4B26.78 GiB8.44 GiB35.85 GiB0.15 GiB12±8.3%
Mistral-Small-4-119B-2603MoEIQ2_XS119B34.48 GiB0.79 GiB35.85 GiB0.15 GiB49±37%
Qwen3.6-35B-A3B-REAM-192-hereticMoEQ5_K_M27.0B34.58 GiB0.70 GiB35.84 GiB0.16 GiB46±37%
MythoMax-L2-Kimiko-v2-13bQ4_K_S13.0B7.09 GiB28.13 GiB35.81 GiB0.19 GiB12±8.3%
MythoMax-L2-13bI1-Q4_K_S13.0B7.09 GiB28.13 GiB35.81 GiB0.19 GiB12±8.3%
MiniMax-M2.1-REAP-139B-A10BMoEI1-IQ1_S139B26.55 GiB8.72 GiB35.80 GiB0.20 GiB20±37%
m51Lab-MiniMax-M2.7-REAP-139B-A10BMoEI1-IQ1_S139B26.55 GiB8.72 GiB35.80 GiB0.20 GiB20±37%
MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-ThinkingQ8_023.4B23.78 GiB11.39 GiB35.76 GiB0.24 GiB12±8.3%
NSFW_13B_sftQ4_013.3B7.03 GiB28.13 GiB35.75 GiB0.25 GiB12±8.3%
GLM-4.5-Air-REAP-82B-A12BMoEIQ2_XXS81.9B28.63 GiB6.47 GiB35.68 GiB0.32 GiB22±37%
Rombo-LLM-V3.0-Qwen-72bI1-IQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
Qwen2.5-72B-Instruct-abliteratedI1-IQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
Qwen2.5-72B-Instruct-abliterated-v2I1-IQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
HuatuoGPT-o1-72BIQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
MiroThinker-v1.0-72BI1-IQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
EVA-Qwen2.5-72B-v0.2IQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
Qwen2.5-Math-72B-InstructIQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
Qwen2.5-72B-InstructIQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
Malaysian-Qwen2.5-72B-InstructI1-IQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
Qwen2.5-72BI1-IQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
magnum-v4-72bI1-IQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
KAT-Dev-72B-ExpIQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
Homer-v1.0-Qwen2.5-72BIQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
Tower-Plus-72B-ultra-uncensored-hereticI1-IQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
Chronos-Platinum-72BIQ2_XXS72.7B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
UI-TARS-72B-DPOIQ2_XXS73.4B23.74 GiB11.25 GiB35.67 GiB0.33 GiB12±8.3%
OLMo-2-1124-13B-InstructIQ4_XS13.7B6.93 GiB28.13 GiB35.65 GiB0.35 GiB12±8.3%
Darwin-35B-A3B-OpusMoEQ8_036.0B34.38 GiB0.70 GiB35.64 GiB0.36 GiB50±37%
Aurora-Code-1MoEQ8_034.7B34.38 GiB0.70 GiB35.64 GiB0.36 GiB50±37%
grug-35b-v2MoEQ8_035.1B34.38 GiB0.70 GiB35.64 GiB0.36 GiB50±37%
grug-35bMoEQ8_035.1B34.38 GiB0.70 GiB35.64 GiB0.36 GiB50±37%
WorldSim-Opus-3.6-35B-A3BMoEQ8_035.1B34.38 GiB0.70 GiB35.64 GiB0.36 GiB50±37%
Qwen3.6-35B-A3B-AnkoMoEQ8_035.1B34.38 GiB0.70 GiB35.64 GiB0.36 GiB50±37%
KAT-Coder-V2.5-DevMoEQ8_034.7B34.38 GiB0.70 GiB35.64 GiB0.36 GiB50±37%
Ornith-1.0-35BMoEQ8_034.7B34.38 GiB0.70 GiB35.64 GiB0.36 GiB50±37%
Nex-N2-miniMoEQ8_035.1B34.38 GiB0.70 GiB35.64 GiB0.36 GiB50±37%
CodeLlama-70b-Instruct-hfI1-Q2_K69.0B23.71 GiB11.25 GiB35.64 GiB0.36 GiB12±8.3%
CodeLlama-70b-Python-hfI1-Q2_K69.0B23.71 GiB11.25 GiB35.64 GiB0.36 GiB12±8.3%
Nous-Hermes-Llama2-70bI1-Q2_K69.0B23.71 GiB11.25 GiB35.64 GiB0.36 GiB12±8.3%
Midnight-Miqu-70B-v1.5I1-Q2_K69.0B23.71 GiB11.25 GiB35.64 GiB0.36 GiB12±8.3%
KafkaLM-70B-German-V0.1Q2_K69.0B23.71 GiB11.25 GiB35.64 GiB0.36 GiB12±8.3%
WhiteRabbitNeo-13B-v1Q4_K_S13.0B6.91 GiB28.13 GiB35.63 GiB0.37 GiB12±8.3%
Orca-2-13b-Alpaca-UncensoredI1-Q4_K_S13.0B6.91 GiB28.13 GiB35.63 GiB0.37 GiB12±8.3%
WizardLM-13B-UncensoredI1-Q4_K_S13.0B6.91 GiB28.13 GiB35.63 GiB0.37 GiB12±8.3%
WizardCoder-Python-13B-V1.0I1-Q4_K_S13.0B6.91 GiB28.13 GiB35.63 GiB0.37 GiB12±8.3%
Guanaco-13B-UncensoredI1-Q4_K_S13.0B6.91 GiB28.13 GiB35.63 GiB0.37 GiB12±8.3%
Wizard-Vicuna-13B-UncensoredQ4_K_S13.0B6.91 GiB28.13 GiB35.63 GiB0.37 GiB12±8.3%
WizardLM-13b-V1.0-UncensoredQ4_K_S13.0B6.91 GiB28.13 GiB35.63 GiB0.37 GiB12±8.3%
Qwen3.6-35B-A3B-uncensored-hereticMoEQ8_035.1B34.37 GiB0.70 GiB35.63 GiB0.37 GiB50±37%
Qwen35B-Agent-R2-AbliteratedMoEQ8_034.7B34.37 GiB0.70 GiB35.63 GiB0.37 GiB50±37%
spoomplesmaxx-flash-35B-A3MoEQ8_035.1B34.37 GiB0.70 GiB35.63 GiB0.37 GiB50±37%
Holo-3.1-35B-A3BMoEQ8_035.1B34.37 GiB0.70 GiB35.63 GiB0.37 GiB50±37%
Qwen-AgentWorld-35B-A3BMoEQ8_034.7B34.37 GiB0.70 GiB35.63 GiB0.37 GiB50±37%
Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoEQ8_036.0B34.37 GiB0.70 GiB35.63 GiB0.37 GiB50±37%
Qwen3.6-35B-A3B-abliterated-MAXMoEQ8_035.1B34.37 GiB0.70 GiB35.63 GiB0.37 GiB50±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M4 Max run?
1996 of 2118 indexed open-weight models fit a Apple M4 Max at 131,072 context with q4_0 KV cache, the largest being Kimi-Dev-72B at UD-IQ2_XXS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M4 Max actually have?
Its nameplate is 48 GB, but about 33.48 GiB is available to a model once driver and compositor overhead is accounted for, and only 36 GB of the pool can be allocated to the GPU at all.
Is a Apple M4 Max fast for local AI?
Its memory bandwidth is 546 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.