Apple · apple

Apple M3 Pro

Apple M3 Pro has 36 GB of unified memory at 154 GB/s — about 25.11 GiB usable after driver and compositor overhead. 2000 of 2118 indexed models fit at 32K context with q4_0 KV. Note only 27 GB of its 36 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
36 GB
LPDDR5-6400
Bandwidth
154 GB/s
192-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1717vision language 179audio tts 21audio asr 39video 16image 2embedding 26

What fits at 32K context

largest quantization that fits, per model · 2000 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Step-3.5-Flash-REAP-121B-A11BI1-IQ1_S121B22.74 GiB3.67 GiB26.98 GiB0.02 GiB5±8.3%
Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingQ5_K_S39.5B25.49 GiB0.84 GiB26.95 GiB0.05 GiB5±8.3%
Gemma-4-31B-Isometry-RPI1-Q6_K32.7B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
Gemma-4-Dark-Gemistry-31BI1-Q6_K32.7B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
Prosopon-31BI1-Q6_K32.7B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
Gemma-4-Novelist-Eclipse-31BI1-Q6_K32.7B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
Giftige-Blume-31B-v1-StyleSwapI1-Q6_K32.7B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
G4-MeroMero-31B-StyleSwapI1-Q6_K32.7B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
Gemma-4-31B-StyleTune-heretic-araI1-Q6_K32.7B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
Pantheon-Reasoning-31B-1.1I1-Q6_K32.7B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
Gemma-4-31B-StyleTuneI1-Q6_K32.7B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
Barcenas-StyleTune-31B-FableI1-Q6_K32.1B24.55 GiB1.74 GiB26.91 GiB0.09 GiB5±8.3%
WizardLM-Uncensored-SuperCOT-StoryTelling-30bQ2_K32.5B12.58 GiB13.71 GiB26.91 GiB0.09 GiB5±8.3%
Wizard-Vicuna-30B-UncensoredQ2_K32.5B12.58 GiB13.71 GiB26.91 GiB0.09 GiB5±8.3%
Kimi-Linear-48B-A3B-InstructMoEIQ4_NL49.1B26.05 GiB0.27 GiB26.87 GiB0.13 GiB5±8.3%
Qwen3.6-34B-80L-Fable-5-HereticI1-Q6_K33.4B25.54 GiB0.70 GiB26.86 GiB0.14 GiB5±8.3%
gemma-4-26B-A4B-itMoEQ8_026.5B25.89 GiB0.43 GiB26.86 GiB0.14 GiB5±8.3%
Laguna-S-2.1MoEIQ1_M118B25.75 GiB0.46 GiB26.78 GiB0.22 GiB22±37%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedIQ1_M109B24.51 GiB1.69 GiB26.77 GiB0.23 GiB15±37%
Noromaid-20b-v0.1.1I1-Q6_K20.0B15.28 GiB10.90 GiB26.77 GiB0.23 GiB5±8.3%
Nethena-20BQ6_K20.0B15.28 GiB10.90 GiB26.77 GiB0.23 GiB5±8.3%
Huihui-Qwen3-Coder-Next-abliteratedMoEQ2_K79.7B26.00 GiB0.21 GiB26.76 GiB0.24 GiB28±37%
c4ai-command-r-08-2024Q6_K32.3B24.68 GiB1.41 GiB26.75 GiB0.25 GiB5±8.3%
Seed-OSS-36B-Instruct-biprojected-norm-preserving-abliteratedI1-Q5_K_M36.2B23.84 GiB2.25 GiB26.74 GiB0.26 GiB5±8.3%
Seed-OSS-36B-InstructQ5_K_M36.2B23.84 GiB2.25 GiB26.74 GiB0.26 GiB5±8.3%
Hermes-4.3-36B-hereticI1-Q5_K_M36.2B23.84 GiB2.25 GiB26.74 GiB0.26 GiB5±8.3%
Hermes-4.3-36BQ5_K_M36.2B23.84 GiB2.25 GiB26.74 GiB0.26 GiB5±8.3%
Seed-OSS-36B-BaseQ5_K_M36.2B23.84 GiB2.25 GiB26.74 GiB0.26 GiB5±8.3%
archangel_sft-kto_llama30bI1-IQ3_XS32.5B12.40 GiB13.71 GiB26.73 GiB0.27 GiB5±8.3%
Qwen3.6-27B-A3B-CoderMoEQ8_026.7B25.99 GiB0.18 GiB26.72 GiB0.28 GiB21±37%
Goetia-26B-A4B-v1.4MoEQ8_026.0B25.75 GiB0.43 GiB26.72 GiB0.28 GiB5±8.3%
G4-Moonlight-Dusk-26B-A4B-hereticMoEQ8_026.5B25.75 GiB0.43 GiB26.72 GiB0.28 GiB5±8.3%
G4-Moonlight-Dusk-26B-A4BMoEQ8_026.5B25.75 GiB0.43 GiB26.72 GiB0.28 GiB5±8.3%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEQ8_026.5B25.75 GiB0.43 GiB26.72 GiB0.28 GiB5±8.3%
Chimera-X-26B-A4BMoEQ8_026.5B25.75 GiB0.43 GiB26.72 GiB0.28 GiB5±8.3%
Pantheon-Reasoning-26B-A4B-1.1MoEQ8_026.5B25.75 GiB0.43 GiB26.72 GiB0.28 GiB5±8.3%
Gemma-4-26B-A4B-StyleTune-V2MoEQ8_026.5B25.75 GiB0.43 GiB26.72 GiB0.28 GiB5±8.3%
Gemma-4-26B-A4B-StyleTuneMoEQ8_026.5B25.75 GiB0.43 GiB26.72 GiB0.28 GiB5±8.3%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEQ8_025.8B25.75 GiB0.43 GiB26.72 GiB0.28 GiB5±8.3%
EXAONE-4.5-33BI1-Q6_K34.4B25.27 GiB0.80 GiB26.72 GiB0.28 GiB5±8.3%
IQuest-Coder-V1-40B-InstructI1-Q4_139.8B23.24 GiB2.81 GiB26.70 GiB0.30 GiB5±8.3%
command-r-35b-writer-v2I1-IQ3_S35.0B14.77 GiB11.25 GiB26.68 GiB0.32 GiB5±8.3%
Apertus-70B-Instruct-2509UD-IQ2_M70.6B23.12 GiB2.81 GiB26.66 GiB0.34 GiB5±8.3%
Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingI1-Q5_K_S39.5B25.20 GiB0.84 GiB26.65 GiB0.35 GiB5±8.3%
Qwen3.5-40B-RoughHouse-Claude-4.6-Opus-Polar-Deckard-Uncensored-Heretic-ThinkingI1-Q5_K_S39.5B25.20 GiB0.84 GiB26.65 GiB0.35 GiB5±8.3%
Qwen3.5-40B-Claude-4.5-Opus-High-Reasoning-Thinking-uncensored-hereticQ5_K_S39.5B25.20 GiB0.84 GiB26.65 GiB0.35 GiB5±8.3%
Huihui-GLM-4.7-Flash-abliterated-57BMoEI1-Q3_K_M57.3B24.87 GiB1.18 GiB26.64 GiB0.36 GiB16±37%
Hypernova-60B-2605MoEI1-IQ2_XXS58.7B25.80 GiB0.29 GiB26.62 GiB0.38 GiB22±37%
dolphin-2.6-mixtral-8x7bMoEI1-Q4_K_S46.7B24.91 GiB1.13 GiB26.62 GiB0.38 GiB8±37%
Nous-Hermes-2-Mixtral-8x7B-DPOMoEQ4_K_S46.7B24.91 GiB1.13 GiB26.62 GiB0.38 GiB8±37%
Mixtral-8x7B-Instruct-v0.1MoEQ4_K_S46.7B24.91 GiB1.13 GiB26.62 GiB0.38 GiB8±37%
xLAM-8x7b-rMoEQ4_K_S46.7B24.91 GiB1.13 GiB26.62 GiB0.38 GiB8±37%
dolphin-2.5-mixtral-8x7bMoEQ4_K_S46.7B24.91 GiB1.13 GiB26.62 GiB0.38 GiB8±37%
Mixtral-8x7B-v0.1MoEQ4_K_S46.7B24.91 GiB1.13 GiB26.62 GiB0.38 GiB8±37%
granite-3.1-8b-instructQ8_08.2B24.62 GiB1.41 GiB26.60 GiB0.40 GiB5±8.3%
Trinity-MiniMoEQ8_026.1B25.88 GiB0.17 GiB26.60 GiB0.40 GiB20±37%
Delphi-25B-SimpleRL-MathI1-Q5_K_M25.0B16.52 GiB9.41 GiB26.56 GiB0.44 GiB5±8.3%
granite-34b-code-base-8kI1-Q6_K33.7B25.93 GiB0.00 GiB26.55 GiB0.45 GiB5±8.3%
Skyfall-31B-v4.2-hereticI1-Q6_K31.4B23.96 GiB1.90 GiB26.53 GiB0.47 GiB5±8.3%
Skyfall-31B-v4.2I1-Q6_K31.4B23.96 GiB1.90 GiB26.53 GiB0.47 GiB5±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M3 Pro run?
2000 of 2118 indexed open-weight models fit a Apple M3 Pro at 32,768 context with q4_0 KV cache, the largest being Step-3.5-Flash-REAP-121B-A11B at I1-IQ1_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M3 Pro actually have?
Its nameplate is 36 GB, but about 25.11 GiB is available to a model once driver and compositor overhead is accounted for, and only 27 GB of the pool can be allocated to the GPU at all.
Is a Apple M3 Pro fast for local AI?
Its memory bandwidth is 154 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.