Apple · apple

Apple M3 Pro

Apple M3 Pro has 18 GB of unified memory at 154 GB/s — about 12.56 GiB usable after driver and compositor overhead. 1272 of 2118 indexed models fit at 128K context with q8_0 KV. Note only 14 GB of its 18 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
18 GB
LPDDR5-6400
Bandwidth
154 GB/s
192-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1049video 15vision language 122embedding 26audio tts 21audio asr 38image 1

What fits at 128K context

largest quantization that fits, per model · 1272 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
starcoder2-15bKV unresolvedQ3_K_M16.0B7.54 GiB5.31 GiB13.50 GiB0.00 GiB10±8.3%
Qwen3-4B-BaseQ6_K4.0B3.38 GiB9.56 GiB13.50 GiB0.00 GiB10±8.3%
Wan2.2-S2V-14BQ4_K_M16.3B12.91 GiB0.00 GiB13.49 GiB0.01 GiB10±8.3%
Grug-12BQ5_K_L12.0B8.40 GiB4.50 GiB13.49 GiB0.01 GiB10±8.3%
gemma-4-12B-it-Esper4Q5_K_L12.0B8.40 GiB4.50 GiB13.49 GiB0.01 GiB10±8.3%
gemma-4-12B-itQ5_K_L12.0B8.40 GiB4.50 GiB13.49 GiB0.01 GiB10±8.3%
legitus-instruct-v1I1-Q4_08.1B4.38 GiB8.50 GiB13.49 GiB0.01 GiB10±8.3%
Apertus-8B-Instruct-2509I1-Q4_08.1B4.38 GiB8.50 GiB13.49 GiB0.01 GiB10±8.3%
Mistral-7B-v0.3Q4_K_L7.2B4.40 GiB8.50 GiB13.49 GiB0.01 GiB10±8.3%
granite-4.0-7B-A1B-Creative-v0.1MoEF166.7B12.44 GiB0.53 GiB13.48 GiB0.02 GiB26±37%
GRM-2.6-Plus-0628IQ1_M27.8B8.62 GiB4.25 GiB13.48 GiB0.02 GiB10±8.3%
MiniCPM-Llama3-V-2_5IQ4_NL8.5B4.38 GiB8.50 GiB13.47 GiB0.03 GiB10±8.3%
Llama3-ChatQA-1.5-8BIQ4_NL8.0B4.38 GiB8.50 GiB13.47 GiB0.03 GiB10±8.3%
grok-oss-Apollyon-8BIQ4_NL8.0B4.38 GiB8.50 GiB13.47 GiB0.03 GiB10±8.3%
Turkish-Llama-8b-Instruct-v0.1IQ4_NL8.0B4.38 GiB8.50 GiB13.47 GiB0.03 GiB10±8.3%
granite-8b-code-instruct-4kI1-IQ3_S8.1B3.32 GiB9.56 GiB13.47 GiB0.03 GiB10±8.3%
granite-8b-code-base-4kI1-IQ3_S8.1B3.32 GiB9.56 GiB13.47 GiB0.03 GiB10±8.3%
OLMoE-1B-7B-0924-InstructMoEI1-Q5_K_S6.9B4.45 GiB8.50 GiB13.47 GiB0.03 GiB8±37%
Goetia-26B-A4B-v1.4MoEI1-Q2_K_S26.0B10.12 GiB2.81 GiB13.47 GiB0.03 GiB10±8.3%
G4-Moonlight-Dusk-26B-A4B-hereticMoEI1-Q2_K_S26.5B10.12 GiB2.81 GiB13.47 GiB0.03 GiB10±8.3%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEI1-Q2_K_S26.5B10.12 GiB2.81 GiB13.47 GiB0.03 GiB10±8.3%
G4-Moonlight-Dusk-26B-A4BMoEI1-Q2_K_S26.5B10.12 GiB2.81 GiB13.47 GiB0.03 GiB10±8.3%
Chimera-X-26B-A4BMoEI1-Q2_K_S26.5B10.12 GiB2.81 GiB13.47 GiB0.03 GiB10±8.3%
Pantheon-Reasoning-26B-A4B-1.1MoEI1-Q2_K_S26.5B10.12 GiB2.81 GiB13.47 GiB0.03 GiB10±8.3%
Gemma-4-26B-A4B-StyleTune-V2MoEI1-Q2_K_S26.5B10.12 GiB2.81 GiB13.47 GiB0.03 GiB10±8.3%
Gemma-4-26B-A4B-StyleTuneMoEI1-Q2_K_S26.5B10.12 GiB2.81 GiB13.47 GiB0.03 GiB10±8.3%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEI1-Q2_K_S25.8B10.12 GiB2.81 GiB13.47 GiB0.03 GiB10±8.3%
Marco-Mini-InstructMoEI1-Q2_K_S17.3B5.51 GiB7.44 GiB13.47 GiB0.03 GiB9±37%
Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingI1-Q3_K_M21.3B9.67 GiB3.19 GiB13.47 GiB0.03 GiB10±8.3%
Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-ThinkingI1-Q3_K_M21.3B9.67 GiB3.19 GiB13.47 GiB0.03 GiB10±8.3%
gemma-4-12BQ5_K_M12.0B8.37 GiB4.50 GiB13.47 GiB0.03 GiB10±8.3%
Assistant_Pepe_8BQ4_K_S4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
gpt-oss-20b-hereticMoEIQ3_XS20.9B11.32 GiB1.60 GiB13.46 GiB0.04 GiB19±37%
ERNIE-4.5-21B-A3B-ThinkingQ3_K_S21.8B9.17 GiB3.72 GiB13.46 GiB0.04 GiB10±8.3%
ERNIE-4.5-21B-A3B-PTQ3_K_S21.9B9.17 GiB3.72 GiB13.46 GiB0.04 GiB10±8.3%
Qwen3.6-35B-A3B-REAM-160-ru-agentMoEQ3_K_L23.6B11.58 GiB1.33 GiB13.46 GiB0.04 GiB25±37%
Foundation-Sec-8B-InstructI1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Foundation-Sec-8B-Instruct-hereticI1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
dolphin-2.9-llama3-8bQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
saiga_llama3_8bQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Meta-Llama-3-8B-InstructQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Meta-Llama-3-8BQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Llama-3.1-Tulu-3-8BQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Llama-3-Groq-8B-Tool-UseQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
llama3.1-heretic-unsensoredI1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Dolphin3.0-Llama3.1-8BQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
dolphin-2.9.4-llama3.1-8bQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Dolphin3.0-Llama3.1-8B-abliteratedQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
LLAMA-3_8B_Unaligned_BETAQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Llama-3.1-8BQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Anubis-Mini-8B-v1I1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Meta-Llama-3-8BQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Deepseek-R1-Distill-NSFW-RPv1Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
L3.1-Dark-Reasoning-LewdPlay-evo-Hermes-R1-Uncensored-8B-hereticI1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
L3.1-Dark-Reasoning-LewdPlay-evo-Hermes-R1-Uncensored-8BI1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Gluon-8BI1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Llama3.3-8B-Instruct-Thinking-Heretic-Uncensored-Claude-4.5-Opus-High-ReasoningI1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Llama3.3-8B-Instruct-Thinking-Claude-4.5-Opus-High-ReasoningI1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
Hypnos-i1-8BQ4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
grok-oss-Apollyon-8B-hereticI1-Q4_K_S8.0B4.37 GiB8.50 GiB13.46 GiB0.04 GiB10±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing339.31 tok/s305.24343.177
Text generation17.53 tok/s16.9530.517
Benchmarked· n=7

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.

Questions people ask

What AI models can a Apple M3 Pro run?
1272 of 2118 indexed open-weight models fit a Apple M3 Pro at 131,072 context with q8_0 KV cache, the largest being starcoder2-15b at Q3_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M3 Pro actually have?
Its nameplate is 18 GB, but about 12.56 GiB is available to a model once driver and compositor overhead is accounted for, and only 14 GB of the pool can be allocated to the GPU at all.
Is a Apple M3 Pro fast for local AI?
Its memory bandwidth is 154 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.