Apple · apple

Apple M3 Pro

Apple M3 Pro has 36 GB of unified memory at 154 GB/s — about 25.11 GiB usable after driver and compositor overhead. 1999 of 2118 indexed models fit at 16K context with q8_0 KV. Note only 27 GB of its 36 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
36 GB
LPDDR5-6400
Bandwidth
154 GB/s
192-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1716vision language 179audio tts 21audio asr 39video 16image 2embedding 26

What fits at 16K context

largest quantization that fits, per model · 1999 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Assistant_Pepe_70BIQ2_S70.6B23.67 GiB2.66 GiB27.00 GiB0.00 GiB5±8.3%
Noromaid-v0.4-Mixtral-Instruct-8x7b-ZlossMoEQ4_K_S46.7B25.30 GiB1.06 GiB26.95 GiB0.05 GiB8±37%
gemma-4-26B-A4B-itMoEQ8_026.5B25.89 GiB0.49 GiB26.91 GiB0.09 GiB5±8.3%
Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingQ5_K_S39.5B25.49 GiB0.80 GiB26.90 GiB0.10 GiB5±8.3%
Skyfall-31B-v4.2Q6_K31.4B24.43 GiB1.79 GiB26.90 GiB0.10 GiB5±8.3%
EXAONE-4.5-33BI1-Q6_K34.4B25.27 GiB0.98 GiB26.90 GiB0.10 GiB5±8.3%
Kimi-Linear-48B-A3B-InstructMoEIQ4_NL49.1B26.05 GiB0.25 GiB26.86 GiB0.14 GiB5±8.3%
command-r-35b-writer-v2I1-IQ3_M35.0B15.55 GiB10.63 GiB26.83 GiB0.17 GiB5±8.3%
Qwen3.6-34B-80L-Fable-5-HereticI1-Q6_K33.4B25.54 GiB0.66 GiB26.82 GiB0.18 GiB5±8.3%
Laguna-S-2.1MoEIQ1_M118B25.75 GiB0.47 GiB26.79 GiB0.21 GiB22±37%
Goetia-26B-A4B-v1.4MoEQ8_026.0B25.75 GiB0.49 GiB26.77 GiB0.23 GiB5±8.3%
G4-Moonlight-Dusk-26B-A4B-hereticMoEQ8_026.5B25.75 GiB0.49 GiB26.77 GiB0.23 GiB5±8.3%
G4-Moonlight-Dusk-26B-A4BMoEQ8_026.5B25.75 GiB0.49 GiB26.77 GiB0.23 GiB5±8.3%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEQ8_026.5B25.75 GiB0.49 GiB26.77 GiB0.23 GiB5±8.3%
Chimera-X-26B-A4BMoEQ8_026.5B25.75 GiB0.49 GiB26.77 GiB0.23 GiB5±8.3%
Pantheon-Reasoning-26B-A4B-1.1MoEQ8_026.5B25.75 GiB0.49 GiB26.77 GiB0.23 GiB5±8.3%
Gemma-4-26B-A4B-StyleTune-V2MoEQ8_026.5B25.75 GiB0.49 GiB26.77 GiB0.23 GiB5±8.3%
Gemma-4-26B-A4B-StyleTuneMoEQ8_026.5B25.75 GiB0.49 GiB26.77 GiB0.23 GiB5±8.3%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEQ8_025.8B25.75 GiB0.49 GiB26.77 GiB0.23 GiB5±8.3%
Huihui-Qwen3-Coder-Next-abliteratedMoEQ2_K79.7B26.00 GiB0.20 GiB26.74 GiB0.26 GiB28±37%
Qwen3.6-27B-A3B-CoderMoEQ8_026.7B25.99 GiB0.17 GiB26.71 GiB0.29 GiB21±37%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedIQ1_M109B24.51 GiB1.59 GiB26.68 GiB0.32 GiB15±37%
c4ai-command-r-08-2024Q6_K32.3B24.68 GiB1.33 GiB26.68 GiB0.32 GiB5±8.3%
WizardLM-Uncensored-SuperCOT-StoryTelling-30bQ3_K_S32.5B13.10 GiB12.95 GiB26.67 GiB0.33 GiB5±8.3%
Wizard-Vicuna-30B-UncensoredI1-IQ3_S32.5B13.10 GiB12.95 GiB26.67 GiB0.33 GiB5±8.3%
archangel_sft-kto_llama30bI1-IQ3_S32.5B13.10 GiB12.95 GiB26.67 GiB0.33 GiB5±8.3%
Trinity-MiniMoEQ8_026.1B25.88 GiB0.20 GiB26.62 GiB0.38 GiB20±37%
Hypernova-60B-2605MoEI1-IQ2_XXS58.7B25.80 GiB0.28 GiB26.61 GiB0.39 GiB22±37%
Seed-OSS-36B-Instruct-biprojected-norm-preserving-abliteratedI1-Q5_K_M36.2B23.84 GiB2.13 GiB26.61 GiB0.39 GiB5±8.3%
Seed-OSS-36B-InstructQ5_K_M36.2B23.84 GiB2.13 GiB26.61 GiB0.39 GiB5±8.3%
Hermes-4.3-36B-hereticI1-Q5_K_M36.2B23.84 GiB2.13 GiB26.61 GiB0.39 GiB5±8.3%
Hermes-4.3-36BQ5_K_M36.2B23.84 GiB2.13 GiB26.61 GiB0.39 GiB5±8.3%
Seed-OSS-36B-BaseQ5_K_M36.2B23.84 GiB2.13 GiB26.61 GiB0.39 GiB5±8.3%
Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingI1-Q5_K_S39.5B25.20 GiB0.80 GiB26.61 GiB0.39 GiB5±8.3%
Qwen3.5-40B-RoughHouse-Claude-4.6-Opus-Polar-Deckard-Uncensored-Heretic-ThinkingI1-Q5_K_S39.5B25.20 GiB0.80 GiB26.61 GiB0.39 GiB5±8.3%
Qwen3.5-40B-Claude-4.5-Opus-High-Reasoning-Thinking-uncensored-hereticQ5_K_S39.5B25.20 GiB0.80 GiB26.61 GiB0.39 GiB5±8.3%
Huihui-GLM-4.7-Flash-abliterated-57BMoEI1-Q3_K_M57.3B24.87 GiB1.11 GiB26.57 GiB0.43 GiB17±37%
dolphin-2.6-mixtral-8x7bMoEI1-Q4_K_S46.7B24.91 GiB1.06 GiB26.56 GiB0.44 GiB9±37%
Nous-Hermes-2-Mixtral-8x7B-DPOMoEQ4_K_S46.7B24.91 GiB1.06 GiB26.56 GiB0.44 GiB9±37%
Mixtral-8x7B-Instruct-v0.1MoEQ4_K_S46.7B24.91 GiB1.06 GiB26.56 GiB0.44 GiB9±37%
xLAM-8x7b-rMoEQ4_K_S46.7B24.91 GiB1.06 GiB26.56 GiB0.44 GiB9±37%
dolphin-2.5-mixtral-8x7bMoEQ4_K_S46.7B24.91 GiB1.06 GiB26.56 GiB0.44 GiB9±37%
Mixtral-8x7B-v0.1MoEQ4_K_S46.7B24.91 GiB1.06 GiB26.56 GiB0.44 GiB9±37%
granite-34b-code-base-8kI1-Q6_K33.7B25.93 GiB0.00 GiB26.55 GiB0.45 GiB5±8.3%
IQuest-Coder-V1-40B-InstructI1-Q4_139.8B23.24 GiB2.66 GiB26.54 GiB0.46 GiB5±8.3%
granite-3.1-8b-instructQ8_08.2B24.62 GiB1.33 GiB26.52 GiB0.48 GiB5±8.3%
Apertus-70B-Instruct-2509UD-IQ2_M70.6B23.12 GiB2.66 GiB26.51 GiB0.49 GiB5±8.3%
Olmo-3.1-32B-InstructQ6_K_L32.2B24.86 GiB0.98 GiB26.49 GiB0.51 GiB5±8.3%
Olmo-3.1-32B-ThinkQ6_K_L32.2B24.86 GiB0.98 GiB26.49 GiB0.51 GiB5±8.3%
Olmo-3-32B-ThinkQ6_K_L32.2B24.86 GiB0.98 GiB26.49 GiB0.51 GiB5±8.3%
Qwen3.5-122B-A10B-hereticMoEI1-IQ1_M123B25.70 GiB0.20 GiB26.48 GiB0.52 GiB26±37%
GLM-Z1-32B-0414Q6_K_L32.6B25.31 GiB0.51 GiB26.46 GiB0.54 GiB5±8.3%
GLM-4-32B-0414Q6_K_L32.6B25.31 GiB0.51 GiB26.46 GiB0.54 GiB5±8.3%
Qwen3-42B-A3B-2507-Thinking-Abliterated-uncensored-TOTAL-RECALL-v2-Medium-MASTER-CODERMoEI1-Q4_142.4B24.78 GiB1.11 GiB26.44 GiB0.56 GiB17±37%
Skyfall-31B-v4.2-hereticI1-Q6_K31.4B23.96 GiB1.79 GiB26.42 GiB0.58 GiB5±8.3%
Ace-Step1.5Q5_K_M160M25.46 GiB0.42 GiB26.42 GiB0.58 GiB5±8.3%
c4ai-command-r-plus-08-2024IQ1_M104B23.49 GiB2.13 GiB26.34 GiB0.66 GiB5±8.3%
EXAONE-4.0-32BQ6_K_L32.0B24.69 GiB0.98 GiB26.32 GiB0.68 GiB5±8.3%
Open_Gpt4_8x7B_v0.1MoEQ4_046.7B24.63 GiB1.06 GiB26.28 GiB0.72 GiB9±37%
dolphin-2.7-mixtral-8x7bMoEQ4_046.7B24.63 GiB1.06 GiB26.27 GiB0.73 GiB9±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M3 Pro run?
1999 of 2118 indexed open-weight models fit a Apple M3 Pro at 16,384 context with q8_0 KV cache, the largest being Assistant_Pepe_70B at IQ2_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M3 Pro actually have?
Its nameplate is 36 GB, but about 25.11 GiB is available to a model once driver and compositor overhead is accounted for, and only 27 GB of the pool can be allocated to the GPU at all.
Is a Apple M3 Pro fast for local AI?
Its memory bandwidth is 154 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.