Apple · apple

Apple M3 Pro

Apple M3 Pro has 36 GB of unified memory at 154 GB/s — about 25.11 GiB usable after driver and compositor overhead. 1835 of 2118 indexed models fit at 128K context with q8_0 KV. Note only 27 GB of its 36 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
36 GB
LPDDR5-6400
Bandwidth
154 GB/s
192-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1560vision language 172audio asr 39video 16image 1embedding 26audio tts 21

What fits at 128K context

largest quantization that fits, per model · 1835 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
WizardCoder-Python-34B-V1.0I1-Q3_K_S33.7B13.60 GiB12.75 GiB27.00 GiB0.00 GiB5±8.3%
Phind-CodeLlama-34B-Python-v1I1-Q3_K_S33.7B13.60 GiB12.75 GiB27.00 GiB0.00 GiB5±8.3%
Phind-CodeLlama-34B-v2I1-Q3_K_S33.7B13.60 GiB12.75 GiB27.00 GiB0.00 GiB5±8.3%
CodeLlama-34b-instruct-hfQ3_K_S33.7B13.60 GiB12.75 GiB27.00 GiB0.00 GiB5±8.3%
WizardLM-1.0-Uncensored-CodeLlama-34bQ3_K_S33.7B13.60 GiB12.75 GiB27.00 GiB0.00 GiB5±8.3%
GLM-4.7-Flash-DerestrictedMoEI1-Q6_K31.2B22.92 GiB3.51 GiB26.99 GiB0.01 GiB11±37%
Huihui-GLM-4.7-Flash-abliteratedMoEI1-Q6_K31.2B22.92 GiB3.51 GiB26.99 GiB0.01 GiB11±37%
GLM-4.7-Flash-Claude-Opus-4.5-High-Reasoning-DistillMoEQ6_K31.2B22.92 GiB3.51 GiB26.99 GiB0.01 GiB11±37%
Laguna-S-2.1MoEIQ1_S118B23.15 GiB3.26 GiB26.98 GiB0.02 GiB13±37%
spoomplesmaxx-v2.1-30BI1-Q2_K_S28.9B9.31 GiB17.00 GiB26.97 GiB0.03 GiB5±8.3%
Huihui-granite-4.1-30b-abliteratedI1-Q2_K_S28.9B9.31 GiB17.00 GiB26.97 GiB0.03 GiB5±8.3%
granite-4.1-30b-hereticI1-Q2_K_S28.9B9.31 GiB17.00 GiB26.97 GiB0.03 GiB5±8.3%
ALIA-40b-fc-2606I1-IQ2_M40.4B13.54 GiB12.75 GiB26.96 GiB0.04 GiB5±8.3%
ALIA-40b-instruct-2606I1-IQ2_M40.4B13.54 GiB12.75 GiB26.96 GiB0.04 GiB5±8.3%
Qwen3-Coder-NextMoEUD-IQ1_S79.7B20.03 GiB6.38 GiB26.95 GiB0.05 GiB9±37%
Phi-3.5-mini-instructIQ1_M3.8B0.88 GiB25.50 GiB26.94 GiB0.06 GiB5±8.3%
Qwen3.5-99BMoEI1-IQ2_XXS99.0B24.77 GiB1.59 GiB26.94 GiB0.06 GiB17±37%
Olmo-3.1-32B-InstructQ5_K_L32.2B21.58 GiB4.70 GiB26.93 GiB0.07 GiB5±8.3%
Olmo-3.1-32B-ThinkQ5_K_L32.2B21.58 GiB4.70 GiB26.93 GiB0.07 GiB5±8.3%
Olmo-3-32B-ThinkQ5_K_L32.2B21.58 GiB4.70 GiB26.93 GiB0.07 GiB5±8.3%
CallerIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Dumpling-Qwen2.5-32BIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
OREAL-32BIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
QwQ-32B-Preview-abliterated-linear25I1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
openhands-lm-32b-v0.1I1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-Coder-32B-abliteratedI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
m1-32bI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
XMainframe-v2-Instruct-32bI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-Coder-32B-Python-SpecialistI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-32b-RP-InkI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
LongWriter-Zero-32BIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
OpenCodeReasoning-Nemotron-32B-IOIIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-Coder-32B-Instruct-abliteratedIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
OlympicCoder-32BIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
OpenCodeReasoning-Nemotron-32BIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
OpenThinker-32BIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
QwQ-32B-ArliAI-RpR-v4IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-Coder-32B-InstructIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-Coder-32BIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
QwQ-32B-abliteratedIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
DeepSeek-R1-Distill-Qwen-32B-hereticI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
InnoSpark-HPC-RM-32BI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
OpenThinker2-32BIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
INTELLECT-2IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-32B-InstructIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-Coder-32B-Instruct-UncensoredI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
QwQ-32B-PreviewIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
TinyR1-32B-PreviewIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
deepseek-r1-qwen-2.5-32B-ablatedIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Rombos-LLM-V2.5-Qwen-32bIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
DeepSeek-R1-Distill-Qwen-32B-abliteratedIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-32B-ArliAI-RPMax-v1.3IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
DeepSeek-R1-Distill-Qwen-32BIQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
Qwen2.5-VL-32B-InstructIQ2_XS33.5B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
EVA-Qwen2.5-32B-v0.2IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
EVA-Qwen2.5-32B-v0.1IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
cogito-v1-preview-qwen-32BI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
QwQ-32B-Snowdrop-v0I1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
DeepSeek-R1-Distill-Qwen-32B-UncensoredI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
RoguePlanet-DeepSeek-R1-Qwen-32B-RPI1-IQ2_XS32.8B9.27 GiB17.00 GiB26.92 GiB0.08 GiB5±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M3 Pro run?
1835 of 2118 indexed open-weight models fit a Apple M3 Pro at 131,072 context with q8_0 KV cache, the largest being WizardCoder-Python-34B-V1.0 at I1-Q3_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M3 Pro actually have?
Its nameplate is 36 GB, but about 25.11 GiB is available to a model once driver and compositor overhead is accounted for, and only 27 GB of the pool can be allocated to the GPU at all.
Is a Apple M3 Pro fast for local AI?
Its memory bandwidth is 154 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.