Apple · apple

Apple M3 Max

Apple M3 Max has 96 GB of unified memory at 307 GB/s — about 66.96 GiB usable after driver and compositor overhead. 2071 of 2118 indexed models fit at 32K context with q4_0 KV. Note only 72 GB of its 96 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
96 GB
LPDDR5-6400
Bandwidth
307 GB/s
384-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1781vision language 186image 2audio asr 39audio tts 21video 16embedding 26

What fits at 32K context

largest quantization that fits, per model · 2071 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Behemoth-X-123B-v2Q4_K_M123B68.19 GiB3.09 GiB71.99 GiB0.01 GiB4±8.3%
Mistral-Large-Instruct-2411Q4_K_M123B68.19 GiB3.09 GiB71.99 GiB0.01 GiB4±8.3%
MiniMax-M2.1-REAP-139B-A10BMoEI1-IQ4_XS139B69.15 GiB2.18 GiB71.86 GiB0.14 GiB14±37%
m51Lab-MiniMax-M2.7-REAP-139B-A10BMoEI1-IQ4_XS139B69.15 GiB2.18 GiB71.86 GiB0.14 GiB14±37%
Mistral-Medium-3.5-128BQ4_K_S128B68.01 GiB3.09 GiB71.81 GiB0.19 GiB4±8.3%
MiniMax-M2.5MoEUD-IQ2_XXS229B69.03 GiB2.18 GiB71.75 GiB0.25 GiB16±37%
MiniMax-M2.1MoEUD-IQ2_XXS229B68.98 GiB2.18 GiB71.69 GiB0.31 GiB16±37%
MiniMax-M2MoEUD-IQ2_XXS229B68.92 GiB2.18 GiB71.63 GiB0.37 GiB16±37%
c4ai-command-r-plus-08-2024Q5_K_M104B68.57 GiB2.25 GiB71.55 GiB0.45 GiB4±8.3%
OYM-Qimi-122B-A10B-K2.6MoEI1-Q4_K_M125B70.64 GiB0.21 GiB71.42 GiB0.58 GiB20±37%
Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliteratedMoEQ4_K_M123B70.63 GiB0.21 GiB71.42 GiB0.58 GiB20±37%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedQ5_K_S109B69.16 GiB1.69 GiB71.42 GiB0.58 GiB14±37%
GLM-4.5-Air-DerestrictedMoEQ4_K_L110B68.88 GiB1.62 GiB71.08 GiB0.92 GiB14±37%
Step-3.7-FlashQ2_K_L201B66.82 GiB3.67 GiB71.06 GiB0.94 GiB4±8.3%
dots.llm1.instMoEIQ3_XS143B61.71 GiB8.72 GiB71.00 GiB1.00 GiB9±37%
GLM-4.7-REAP-218B-A32BMoEUD-IQ2_XXS218B67.13 GiB3.23 GiB70.96 GiB1.04 GiB11±37%
GLM-4.5-AirMoEQ4_K_M110B68.45 GiB1.62 GiB70.65 GiB1.35 GiB14±37%
Qwen3.5-122B-A10BMoEQ4_K_S125B69.66 GiB0.21 GiB70.45 GiB1.55 GiB21±37%
Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16MoEQ8_035.1B69.57 GiB0.18 GiB70.30 GiB1.70 GiB21±37%
Seed-OSS-36B-InstructBF1636.2B67.35 GiB2.25 GiB70.24 GiB1.76 GiB4±8.3%
Hermes-4.3-36BBF1636.2B67.35 GiB2.25 GiB70.24 GiB1.76 GiB4±8.3%
Mixtral-8x22B-Instruct-v0.1MoEQ3_K_L141B67.60 GiB1.97 GiB70.18 GiB1.82 GiB7±37%
Mixtral-8x22B-v0.1MoEQ3_K_L141B67.60 GiB1.97 GiB70.17 GiB1.83 GiB7±37%
Mixtral-8x22B-v0.1MoEQ3_K_L141B67.60 GiB1.97 GiB70.17 GiB1.83 GiB7±37%
Devstral-2-123B-Instruct-2512Q4_K_S125B66.36 GiB3.09 GiB70.16 GiB1.84 GiB4±8.3%
XORTRON-NXTXPRTXXLI1-Q4_K_S128B66.36 GiB3.09 GiB70.16 GiB1.84 GiB4±8.3%
MiMo-V2-FlashMoEKV unresolvedIQ2_XXS310B68.47 GiB1.05 GiB70.13 GiB1.87 GiB19±37%
Laguna-S-2.1MoEQ4_1118B68.96 GiB0.46 GiB69.99 GiB2.01 GiB19±37%
Qwen3.5-122B-A10B-hereticMoEI1-Q4_K_M123B69.11 GiB0.21 GiB69.90 GiB2.10 GiB21±37%
HunyuanImage-2.1Q6_K17.5B68.97 GiB0.00 GiB69.57 GiB2.43 GiB4±8.3%
step-3.5-flashQ2_K_L199B65.26 GiB3.67 GiB69.50 GiB2.50 GiB4±8.3%
Mistral-Small-4-119B-2603MoEUD-Q4_K_M119B68.70 GiB0.20 GiB69.48 GiB2.52 GiB21±37%
gpt-oss-120b-Uncensored-xCloudMoEI1-Q4_1117B68.42 GiB0.32 GiB69.28 GiB2.72 GiB21±37%
gpt-oss-120b-abliteratedMoEI1-Q4_1117B68.42 GiB0.32 GiB69.28 GiB2.72 GiB21±37%
GLM-4.6VMoEQ4_K_L108B66.89 GiB1.62 GiB69.08 GiB2.92 GiB14±37%
MiMo-V2.5MoEKV unresolvedIQ1_M311B67.01 GiB1.05 GiB68.66 GiB3.34 GiB19±37%
gpt-oss-20b-hereticMoEIQ4_NL20.9B67.58 GiB0.22 GiB68.33 GiB3.67 GiB12±37%
Step-3.5-Flash-REAP-121B-A11BI1-Q4_K_S121B63.83 GiB3.67 GiB68.07 GiB3.93 GiB4±8.3%
MiniMax-M2.7MoEUD-IQ2_M229B65.32 GiB2.18 GiB68.04 GiB3.96 GiB17±37%
CalmeRys-78B-Orpo-v0.1Q6_K78.0B64.27 GiB3.02 GiB67.98 GiB4.02 GiB4±8.3%
calme-2.3-rys-78bQ6_K78.0B64.27 GiB3.02 GiB67.98 GiB4.02 GiB4±8.3%
Qwen3.5-88BMoEI1-Q6_K87.7B67.08 GiB0.21 GiB67.86 GiB4.14 GiB19±37%
GLM-4.5VMoEI1-Q4_K_M108B65.61 GiB1.62 GiB67.80 GiB4.20 GiB15±37%
Qwen2.5-Coder-32B-InstructQ8_032.8B64.86 GiB2.25 GiB67.76 GiB4.24 GiB4±8.3%
Qwen3-235B-A22B-abliteratedMoEI1-IQ2_S235B65.40 GiB1.65 GiB67.64 GiB4.36 GiB15±37%
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEQ4_0124B66.12 GiB0.77 GiB67.44 GiB4.56 GiB18±37%
Qwen3-VL-235B-A22B-ThinkingMoEUD-IQ1_M236B64.90 GiB1.65 GiB67.14 GiB4.86 GiB15±37%
Qwen3-VL-235B-A22B-InstructMoEUD-IQ1_M236B64.83 GiB1.65 GiB67.06 GiB4.94 GiB15±37%
Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-PreservedMoEBF1635.1B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Qwen-AgentWorld-35B-A3BMoEBF1634.7B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Qwable-v1MoEBF1636.0B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Salience-1.5-ProMoEBF1636.0B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
T-SearchMoEBF1636.0B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Qwen35B-Agent-R2MoEF1634.7B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Ornith-1.0-35B-Heretic-MTPMoEBF1666.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Ornith-1.0-35BMoEBF1634.7B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Carnice-Qwen3.6-MoE-35B-A3BMoEF1636.0B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoEF1636.0B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Qwen3.6-35B-A3BMoEBF1636.0B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-DistilledMoEF1636.0B66.19 GiB0.18 GiB66.92 GiB5.08 GiB22±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M3 Max run?
2071 of 2118 indexed open-weight models fit a Apple M3 Max at 32,768 context with q4_0 KV cache, the largest being Behemoth-X-123B-v2 at Q4_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M3 Max actually have?
Its nameplate is 96 GB, but about 66.96 GiB is available to a model once driver and compositor overhead is accounted for, and only 72 GB of the pool can be allocated to the GPU at all.
Is a Apple M3 Max fast for local AI?
Its memory bandwidth is 307 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.