Apple · apple

Apple M3 Ultra

Apple M3 Ultra has 512 GB of unified memory at 819 GB/s — about 357.12 GiB usable after driver and compositor overhead. 2116 of 2118 indexed models fit at 128K context with q8_0 KV. Note only 384 GB of its 512 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
512 GB
LPDDR5-6400
Bandwidth
819 GB/s
1024-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1817vision language 195image 2audio tts 21audio asr 39video 16embedding 26

What fits at 128K context

largest quantization that fits, per model · 2116 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
DeepSeek-V3.2MoEQ4_K_M685B377.56 GiB4.56 GiB382.74 GiB1.26 GiB10±37%
cogito-671b-v2.1MoEQ4_K_M671B377.55 GiB4.56 GiB382.74 GiB1.26 GiB10±37%
Kimi-K2.5MoEIQ3_S1059B377.51 GiB4.56 GiB382.70 GiB1.30 GiB11±37%
DeepSeek-TNG-R1T2-ChimeraMoEQ4_K_M685B377.13 GiB4.56 GiB382.32 GiB1.68 GiB10±37%
cogito-v2-preview-deepseek-671B-MoEMoEQ4_K_M671B377.13 GiB4.56 GiB382.31 GiB1.69 GiB10±37%
GLM-5.1MoEIQ4_XS754B375.69 GiB5.83 GiB382.12 GiB1.88 GiB10±37%
DeepSeek-V3-0324MoEQ4_K_M685B376.89 GiB4.56 GiB382.08 GiB1.92 GiB10±37%
DeepSeek-Prover-V2-671BMoEQ4_K_M685B376.71 GiB4.56 GiB381.90 GiB2.10 GiB10±37%
r1-1776MoEQ4_K_M671B376.65 GiB4.56 GiB381.84 GiB2.16 GiB10±37%
DeepSeek-R1MoEQ4_K_M685B376.65 GiB4.56 GiB381.84 GiB2.16 GiB10±37%
GLM-5MoEIQ4_XS754B375.23 GiB5.83 GiB381.65 GiB2.35 GiB10±37%
GLM-4.5MoEQ8_0358B354.80 GiB24.44 GiB379.83 GiB4.17 GiB6±37%
GLM-4.7MoEQ8_0358B354.80 GiB24.44 GiB379.83 GiB4.17 GiB6±37%
GLM-4.6MoEQ8_0357B353.27 GiB24.44 GiB378.30 GiB5.70 GiB6±37%
GLM-4.6-Derestricted-v3MoEQ8_0357B353.27 GiB24.44 GiB378.30 GiB5.70 GiB6±37%
Qwen3-Coder-REAP-363B-A35BMoEQ8_0363B359.45 GiB16.47 GiB376.50 GiB7.50 GiB6±37%
MiMo-V2.5-ProMoEKV unresolvedUD-IQ3_S1023B351.94 GiB23.24 GiB375.80 GiB8.20 GiB8±37%
Kimi-K2-ThinkingMoEIQ3_XXS1058B367.09 GiB4.56 GiB372.27 GiB11.73 GiB11±37%
DeepSeek-R1-0528MoEQ4_K_S685B367.08 GiB4.56 GiB372.27 GiB11.73 GiB10±37%
DeepSeek-V3.1-TerminusMoEQ4_K_S685B367.08 GiB4.56 GiB372.27 GiB11.73 GiB10±37%
DeepSeek-V3.1MoEQ4_K_S685B367.08 GiB4.56 GiB372.27 GiB11.73 GiB10±37%
NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16MoEUD-Q5_K_S561B364.80 GiB0.00 GiB365.38 GiB18.62 GiB10±37%
Kimi-K2.7-CodeMoEUD-IQ3_XXS1059B351.00 GiB4.56 GiB356.18 GiB27.82 GiB12±37%
GLM-5.2MoEUD-IQ4_NL753B347.07 GiB5.83 GiB353.49 GiB30.51 GiB10±37%
Kimi-K2-InstructMoEQ2_K_L1026B347.81 GiB4.56 GiB352.99 GiB31.01 GiB12±37%
MiniMax-M3MoEQ6_K427B344.03 GiB7.97 GiB352.56 GiB31.44 GiB10±37%
Hermes-3-Llama-3.1-405BQ6_K406B310.08 GiB33.47 GiB344.38 GiB39.62 GiB2±8.3%
Hermes-4-405BQ6_K406B310.08 GiB33.47 GiB344.38 GiB39.62 GiB2±8.3%
Kimi-K2.6MoEQ2_K_L1059B334.66 GiB4.56 GiB339.85 GiB44.15 GiB12±37%
Qwen3-Coder-480B-A35B-InstructMoEQ5_K_M480B317.25 GiB16.47 GiB334.30 GiB49.70 GiB8±37%
dots.llm1.instMoEBF16143B266.00 GiB65.88 GiB332.46 GiB51.54 GiB4±37%
Qwen3.5-397B-A17BMoEQ6_K_L403B325.38 GiB1.99 GiB327.97 GiB56.03 GiB13±37%
Trinity-Large-ThinkingMoEQ6_K_L399B319.94 GiB4.40 GiB324.92 GiB59.08 GiB13±37%
Ornith-1.0-397BMoEQ6_K397B318.87 GiB1.99 GiB321.47 GiB62.53 GiB14±37%
Nex-N2-ProMoEQ6_K397B318.87 GiB1.99 GiB321.47 GiB62.53 GiB14±37%
Llama-4-Maverick-17B-128E-InstructMoEKV unresolvedQ6_K402B306.19 GiB12.75 GiB319.52 GiB64.48 GiB11±37%
Hy3MoEQ8_0299B295.84 GiB21.25 GiB317.67 GiB66.33 GiB8±37%
MiMo-V2.5MoEKV unresolvedQ8_0311B306.67 GiB7.97 GiB315.23 GiB68.77 GiB11±37%
MiMo-V2-FlashMoEKV unresolvedQ8_0310B305.69 GiB7.97 GiB314.25 GiB69.75 GiB11±37%
ERNIE-4.5-300B-A47B-PTQ8_0300B296.43 GiB14.34 GiB311.45 GiB72.55 GiB2±8.3%
Trinity-Large-PreviewMoEQ6_K_L399B305.35 GiB4.40 GiB310.33 GiB73.67 GiB14±37%
Trinity-Large-TrueBaseMoEQ6_K399B305.07 GiB4.40 GiB310.05 GiB73.95 GiB14±37%
GLM-4.6-REAP-268B-A32BMoEQ8_0269B266.14 GiB24.44 GiB291.17 GiB92.83 GiB7±37%
Athene-70BF3270.6B262.84 GiB21.25 GiB284.77 GiB99.23 GiB2±8.3%
grok-2MoEQ8_0270B266.72 GiB17.00 GiB284.41 GiB99.59 GiB4±37%
InklingMoEUD-IQ1_M952B265.46 GiB17.53 GiB283.56 GiB100.44 GiB10±37%
Llama-3_1-Nemotron-51B-InstructF1651.5B95.94 GiB170.00 GiB266.63 GiB117.37 GiB3±8.3%
Llama-3_3-Nemotron-Super-49B-v1_5BF1649.9B92.89 GiB170.00 GiB263.59 GiB120.41 GiB3±8.3%
Valkyrie-49B-v2.1BF1649.9B92.89 GiB170.00 GiB263.59 GiB120.41 GiB3±8.3%
Llama-3_3-Nemotron-Super-49B-v1BF1649.9B92.89 GiB170.00 GiB263.59 GiB120.41 GiB3±8.3%
Solar-Open2-250BMoEQ8_0250B247.86 GiB12.75 GiB261.19 GiB122.81 GiB11±37%
Devstral-2-123B-Instruct-2512BF16125B232.89 GiB23.38 GiB256.97 GiB127.03 GiB3±8.3%
Mistral-Medium-3.5-128BBF16128B232.89 GiB23.38 GiB256.97 GiB127.03 GiB3±8.3%
Qwen3-VL-235B-A22B-InstructMoEQ8_0236B232.77 GiB12.48 GiB245.84 GiB138.16 GiB9±37%
Qwen3-VL-235B-A22B-ThinkingMoEQ8_0236B232.77 GiB12.48 GiB245.84 GiB138.16 GiB9±37%
Qwen3-235B-A22BMoEQ8_0235B232.77 GiB12.48 GiB245.84 GiB138.16 GiB9±37%
Qwen3-235B-A22B-Thinking-2507MoEQ8_0235B232.77 GiB12.48 GiB245.84 GiB138.16 GiB9±37%
Qwen3-235B-A22B-Instruct-2507MoEQ8_0235B232.77 GiB12.48 GiB245.84 GiB138.16 GiB9±37%
MiniMax-M2.1MoEQ8_0229B226.44 GiB16.47 GiB243.44 GiB140.56 GiB10±37%
MiniMax-M2.5MoEQ8_0229B226.44 GiB16.47 GiB243.44 GiB140.56 GiB10±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M3 Ultra run?
2116 of 2118 indexed open-weight models fit a Apple M3 Ultra at 131,072 context with q8_0 KV cache, the largest being DeepSeek-V3.2 at Q4_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M3 Ultra actually have?
Its nameplate is 512 GB, but about 357.12 GiB is available to a model once driver and compositor overhead is accounted for, and only 384 GB of the pool can be allocated to the GPU at all.
Is a Apple M3 Ultra fast for local AI?
Its memory bandwidth is 819 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.