Apple M4 Pro
Apple M4 Pro has 24 GB of unified memory at 273 GB/s — about 16.74 GiB usable after driver and compositor overhead. 1775 of 2118 indexed models fit at 128K context with q4_0 KV. Note only 18 GB of its 24 GB is allocatable to the GPU.
What fits at 128K context
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| internlm2-math-plus-20b | I1-Q4_K_S | 19.9B | 10.62 GiB | 6.75 GiB | 17.98 GiB | 0.02 GiB | 13±8.3% |
| Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved | Q4_K_S | 27.4B | 15.11 GiB | 2.25 GiB | 17.97 GiB | 0.03 GiB | 13±8.3% |
| Qwen3.5-27B-uncensored-heretic-v2-Native-MTP-Preserved | Q4_K_S | 27.4B | 15.11 GiB | 2.25 GiB | 17.97 GiB | 0.03 GiB | 13±8.3% |
| Wan2.1-VACE-14B | Q8_0 | 17.3B | 17.38 GiB | 0.00 GiB | 17.97 GiB | 0.03 GiB | 13±8.3% |
| Phi-3-mini-4k-instructKV unresolved | IQ2_XS | 3.8B | 3.91 GiB | 13.50 GiB | 17.97 GiB | 0.03 GiB | 13±8.3% |
| SOLAR-10.7B-Instruct-v1.0-uncensored | Q8_0 | 10.7B | 10.62 GiB | 6.75 GiB | 17.96 GiB | 0.04 GiB | 13±8.3% |
| Nous-Hermes-2-SOLAR-10.7B | Q8_0 | 10.7B | 10.62 GiB | 6.75 GiB | 17.96 GiB | 0.04 GiB | 13±8.3% |
| SOLAR-10.7B-Instruct-v1.0 | Q8_0 | 10.7B | 10.62 GiB | 6.75 GiB | 17.96 GiB | 0.04 GiB | 13±8.3% |
| InternVL3_5-30B-A3B | Q4_K_M | 30.8B | 17.35 GiB | 0.00 GiB | 17.95 GiB | 0.05 GiB | 13±8.3% |
| reka-flash-3.1 | I1-Q4_K_M | 20.9B | 12.68 GiB | 4.64 GiB | 17.94 GiB | 0.06 GiB | 13±8.3% |
| reka-flash-3 | Q4_K_M | 20.9B | 12.68 GiB | 4.64 GiB | 17.94 GiB | 0.06 GiB | 13±8.3% |
| Snowpiercer-15B-v4 | Q5_K_L | 15.0B | 10.31 GiB | 7.03 GiB | 17.94 GiB | 0.06 GiB | 13±8.3% |
| spoomplesmaxx-v2.1-30B | I1-IQ2_S | 28.9B | 8.28 GiB | 9.00 GiB | 17.94 GiB | 0.06 GiB | 13±8.3% |
| Huihui-granite-4.1-30b-abliterated | I1-IQ2_S | 28.9B | 8.28 GiB | 9.00 GiB | 17.94 GiB | 0.06 GiB | 13±8.3% |
| granite-4.1-30b-heretic | I1-IQ2_S | 28.9B | 8.28 GiB | 9.00 GiB | 17.94 GiB | 0.06 GiB | 13±8.3% |
| UncensoredLM-DeepSeek-R1-Distill-Qwen-14B | Q6_K | 14.2B | 10.87 GiB | 6.47 GiB | 17.94 GiB | 0.06 GiB | 13±8.3% |
| Mistral-MOE-4X7B-Dark-MultiVerse-Uncensored-Enhanced32-24BMoE | Q4_K_S | 24.2B | 12.84 GiB | 4.50 GiB | 17.93 GiB | 0.07 GiB | 7±37% |
| EuroLLM-22B-Instruct-2512 | IQ3_M | 22.6B | 9.72 GiB | 7.59 GiB | 17.93 GiB | 0.07 GiB | 13±8.3% |
| OLMoE-1B-7B-0924-InstructMoE | F16 | 6.9B | 12.89 GiB | 4.50 GiB | 17.91 GiB | 0.09 GiB | 18±37% |
| Magistry-24B-v1.1 | Q3_K_L | 23.6B | 11.60 GiB | 5.63 GiB | 17.90 GiB | 0.10 GiB | 13±8.3% |
| gemma-4-26B-A4B-itMoE | Q4_K_M | 26.5B | 15.87 GiB | 1.49 GiB | 17.89 GiB | 0.11 GiB | 13±8.3% |
| NVIDIA-Nemotron-Nano-12B-v2 | Q5_K_L | 12.3B | 8.55 GiB | 8.72 GiB | 17.89 GiB | 0.11 GiB | 13±8.3% |
| OmniAtlas-Qwen3-30B-A3B | I1-Q4_K_M | 31.7B | 17.28 GiB | 0.00 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| Qwen3-Omni-30B-A3B-Instruct | Q4_K_M | 35.3B | 17.28 GiB | 0.00 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| Qwen3-Omni-30B-A3B-Captioner | I1-Q4_K_M | 31.7B | 17.28 GiB | 0.00 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| Qwen3-Omni-30B-A3B-Thinking | Q4_K_M | 31.7B | 17.28 GiB | 0.00 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| Qwen3.8-27B | Q4_K_S | 27.8B | 15.01 GiB | 2.25 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| Qwen3.6-27B | Q4_K_S | 27.8B | 15.01 GiB | 2.25 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| NousCoder-14B | Q6_K_L | 14.8B | 11.64 GiB | 5.63 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| Qwen3-14B-abliterated | Q6_K_L | 14.8B | 11.64 GiB | 5.63 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| Josiefied-Qwen3-14B-abliterated-v3 | Q6_K_L | 14.8B | 11.64 GiB | 5.63 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| Hermes-4-14B | Q6_K_L | 14.8B | 11.64 GiB | 5.63 GiB | 17.88 GiB | 0.12 GiB | 13±8.3% |
| Qwen3.6-34B-80L-Fable-5-Heretic | I1-IQ3_M | 33.4B | 14.45 GiB | 2.81 GiB | 17.87 GiB | 0.13 GiB | 13±8.3% |
| Phi-3.5-MoE-instructMoEKV unresolved | IQ2_M | 41.9B | 12.82 GiB | 4.50 GiB | 17.87 GiB | 0.13 GiB | 18±37% |
| Qwen3.6-27B-Heretic2-Uncensored-Finetune-Thinking | IQ4_NL | 27.4B | 15.01 GiB | 2.25 GiB | 17.87 GiB | 0.13 GiB | 13±8.3% |
| Le-Chaton-Slim-23BMoE | I1-Q4_1 | 23.3B | 13.65 GiB | 3.66 GiB | 17.87 GiB | 0.13 GiB | 17±37% |
| grug-27b | Q4_0 | 27.4B | 15.00 GiB | 2.25 GiB | 17.87 GiB | 0.13 GiB | 13±8.3% |
| Carnice-V2-27b | Q4_0 | 27.4B | 15.00 GiB | 2.25 GiB | 17.87 GiB | 0.13 GiB | 13±8.3% |
| Fara1.5-27B | Q4_0 | 27.4B | 15.00 GiB | 2.25 GiB | 17.87 GiB | 0.13 GiB | 13±8.3% |
| Trinity-2-Codestral-22B-v0.2 | IQ3_M | 22.2B | 9.37 GiB | 7.88 GiB | 17.86 GiB | 0.14 GiB | 13±8.3% |
| Cydonia-v1.3-Magnum-v4-22B | I1-IQ3_M | 22.2B | 9.37 GiB | 7.88 GiB | 17.86 GiB | 0.14 GiB | 13±8.3% |
| Mistral-Small-22B-ArliAI-RPMax-v1.1 | I1-IQ3_M | 22.2B | 9.37 GiB | 7.88 GiB | 17.86 GiB | 0.14 GiB | 13±8.3% |
| Mistral-Small-Drummer-22B | IQ3_M | 22.2B | 9.37 GiB | 7.88 GiB | 17.86 GiB | 0.14 GiB | 13±8.3% |
| magnum-v4-22b | I1-IQ3_M | 22.2B | 9.37 GiB | 7.88 GiB | 17.86 GiB | 0.14 GiB | 13±8.3% |
| Mistral-Small-Instruct-2409 | IQ3_M | 22.2B | 9.37 GiB | 7.88 GiB | 17.86 GiB | 0.14 GiB | 13±8.3% |
| Codestral-22B-v0.1 | IQ3_M | 22.2B | 9.37 GiB | 7.88 GiB | 17.86 GiB | 0.14 GiB | 13±8.3% |
| Qwen3.5-35B-A3BMoE | UD-IQ4_NL | 36.0B | 16.60 GiB | 0.70 GiB | 17.86 GiB | 0.14 GiB | 45±37% |
| Qwen3.6-28BMoE | I1-Q4_1 | 28.2B | 16.60 GiB | 0.70 GiB | 17.85 GiB | 0.15 GiB | 43±37% |
| Qwen3.5-28BMoE | I1-Q4_1 | 28.7B | 16.60 GiB | 0.70 GiB | 17.85 GiB | 0.15 GiB | 43±37% |
| dolphin-2.9.1-mixtral-1x22bMoE | I1-IQ3_M | 22.2B | 9.37 GiB | 7.88 GiB | 17.85 GiB | 0.15 GiB | 7±37% |
| North-Mini-Code-1.0MoE | Q4_0 | 30.5B | 16.32 GiB | 1.00 GiB | 17.85 GiB | 0.15 GiB | 36±37% |
| Voxtral-Small-24B-2507 | Q3_K_L | 24.3B | 11.55 GiB | 5.63 GiB | 17.84 GiB | 0.16 GiB | 13±8.3% |
| Devstral-Small-2-24B-Instruct-2512 | Q3_K_L | 24.0B | 11.55 GiB | 5.63 GiB | 17.84 GiB | 0.16 GiB | 13±8.3% |
| Transformed-Journey-24B | I1-Q3_K_L | 23.6B | 11.55 GiB | 5.63 GiB | 17.84 GiB | 0.16 GiB | 13±8.3% |
| Mergedonia-AETHER-24B-v1a | I1-Q3_K_L | 23.6B | 11.55 GiB | 5.63 GiB | 17.84 GiB | 0.16 GiB | 13±8.3% |
| Mergedonia-AETHER-24B-v1b | I1-Q3_K_L | 23.6B | 11.55 GiB | 5.63 GiB | 17.84 GiB | 0.16 GiB | 13±8.3% |
| Slimaki-Tavern-24B-v1.3 | I1-Q3_K_L | 23.6B | 11.55 GiB | 5.63 GiB | 17.84 GiB | 0.16 GiB | 13±8.3% |
| Maginum-Cydoms-24B | I1-Q3_K_L | 23.6B | 11.55 GiB | 5.63 GiB | 17.84 GiB | 0.16 GiB | 13±8.3% |
| Maginum-Cydoms-24B-absolute-heresy | I1-Q3_K_L | 23.6B | 11.55 GiB | 5.63 GiB | 17.84 GiB | 0.16 GiB | 13±8.3% |
| Dolphin3.0-R1-Mistral-24B | Q3_K_L | 23.6B | 11.55 GiB | 5.63 GiB | 17.84 GiB | 0.16 GiB | 13±8.3% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Measured on this card
| Workload◍ | Median | Middle 50% | Runs |
|---|---|---|---|
| Prompt processing | 410.46 tok/s | 370.63–447.16 | 6 |
| Text generation | 30.62 tok/s | 20.53–44.90 | 6 |
Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.
Questions people ask
- What AI models can a Apple M4 Pro run?
- 1775 of 2118 indexed open-weight models fit a Apple M4 Pro at 131,072 context with q4_0 KV cache, the largest being internlm2-math-plus-20b at I1-Q4_K_S. That covers text, vision-language, image, video and speech models.
- How much usable memory does a Apple M4 Pro actually have?
- Its nameplate is 24 GB, but about 16.74 GiB is available to a model once driver and compositor overhead is accounted for, and only 18 GB of the pool can be allocated to the GPU at all.
- Is a Apple M4 Pro fast for local AI?
- Its memory bandwidth is 273 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.