Apple · apple
Apple M4
Apple M4 has 16 GB of unified memory at 120 GB/s — about 11.16 GiB usable after driver and compositor overhead. 1712 of 2118 indexed models fit at 64K context with q4_0 KV. Note only 12 GB of its 16 GB is allocatable to the GPU.
Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
LPDDR5X-7500
Bandwidth
120 GB/s
128-bit bus
Tensor FP16
—
dense
TDP
—
vision language 149video 15text 1460image 2audio asr 39embedding 26audio tts 21
What fits at 64K context
largest quantization that fits, per model · 1712 of 2118 indexed
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| Qwen3.5-35B-A3BMoE | IQ2_S | 36.0B | 11.09 GiB | 0.35 GiB | 12.00 GiB | 0.00 GiB | 35±37% |
| Qwen3.6-35B-A3BMoE | IQ2_S | 36.0B | 11.09 GiB | 0.35 GiB | 12.00 GiB | 0.00 GiB | 35±37% |
| Wan2.1-VACE-14B | Q5_K_S | 17.3B | 11.41 GiB | 0.00 GiB | 12.00 GiB | 0.00 GiB | 9±8.3% |
| medgemma-27b-it | I1-Q2_K | 28.8B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| gemma-3-27b-it-abliterated-refined-vision | I1-Q2_K | 27.4B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| gemma-3-27b-it-abliterated | Q2_K | 27.4B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| Nidum-Gemma-3-27B-it-Uncensored | I1-Q2_K | 27.4B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| gemma-3-27b-it | Q2_K | 27.4B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| AtomicGPT-gemma3-27b | I1-Q2_K | 27.4B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| Unbound-v1.12.0-27B | I1-Q2_K | 27.4B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| Mira-v1.12-Ties-27B | I1-Q2_K | 27.4B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| Medgamma27B | I1-Q2_K | 27.0B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| medgemma-27b-text-it | Q2_K | 27.0B | 9.78 GiB | 1.58 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| Qwythos-9B-v2 | Q4_K_M | 9.7B | 10.84 GiB | 0.56 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| GLM-4.7-Flash-REAP-23B-A3BMoE | Q3_K_M | 23.0B | 10.50 GiB | 0.93 GiB | 11.99 GiB | 0.01 GiB | 22±37% |
| Phi-4-reasoning-plus | Q4_K_S | 14.7B | 7.86 GiB | 3.52 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| Phi-4-reasoning | Q4_K_S | 14.7B | 7.86 GiB | 3.52 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| phi-4 | Q4_K_S | 14.7B | 7.86 GiB | 3.52 GiB | 11.99 GiB | 0.01 GiB | 9±8.3% |
| dolphin-2.6-mixtral-8x7bMoE | I1-IQ1_S | 46.7B | 9.15 GiB | 2.25 GiB | 11.98 GiB | 0.02 GiB | 11±37% |
| xLAM-8x7b-rMoE | IQ1_S | 46.7B | 9.15 GiB | 2.25 GiB | 11.98 GiB | 0.02 GiB | 11±37% |
| deepseek-coder-6.7B-kexer | I1-IQ3_XXS | 6.7B | 2.41 GiB | 9.00 GiB | 11.98 GiB | 0.02 GiB | 9±8.3% |
| Magicoder-S-DS-6.7B | I1-IQ3_XXS | 6.7B | 2.41 GiB | 9.00 GiB | 11.98 GiB | 0.02 GiB | 9±8.3% |
| deepseek-coder-6.7b-base | I1-IQ3_XXS | 6.7B | 2.41 GiB | 9.00 GiB | 11.98 GiB | 0.02 GiB | 9±8.3% |
| WizardLM-7B-Uncensored | I1-IQ3_XXS | 6.7B | 2.41 GiB | 9.00 GiB | 11.98 GiB | 0.02 GiB | 9±8.3% |
| Llama-2-7B-32K-Instruct | I1-IQ3_XXS | 6.7B | 2.41 GiB | 9.00 GiB | 11.98 GiB | 0.02 GiB | 9±8.3% |
| Luna-AI-Llama2-Uncensored | I1-IQ3_XXS | 6.7B | 2.41 GiB | 9.00 GiB | 11.98 GiB | 0.02 GiB | 9±8.3% |
| Swallow-7b-NVE-instruct-hf | I1-IQ3_XXS | 6.7B | 2.41 GiB | 9.00 GiB | 11.98 GiB | 0.02 GiB | 9±8.3% |
| Rocinante-XL-16B-v1 | Q3_K_M | 16.1B | 7.59 GiB | 3.80 GiB | 11.98 GiB | 0.02 GiB | 9±8.3% |
| North-Mini-Code-1.0MoE | UD-IQ3_XXS | 30.5B | 10.90 GiB | 0.55 GiB | 11.98 GiB | 0.02 GiB | 27±37% |
| Qwen3-42B-A3B-2507-Thinking-Abliterated-uncensored-TOTAL-RECALL-v2-Medium-MASTER-CODERMoE | I1-IQ1_M | 42.4B | 9.08 GiB | 2.36 GiB | 11.97 GiB | 0.03 GiB | 16±37% |
| InternVL3_5-30B-A3B | IQ3_XXS | 30.8B | 11.38 GiB | 0.00 GiB | 11.97 GiB | 0.03 GiB | 9±8.3% |
| gemma-2-27b-it | IQ2_XS | 27.2B | 7.82 GiB | 3.46 GiB | 11.97 GiB | 0.03 GiB | 9±8.3% |
| magnum-v4-27b | IQ2_XS | 27.2B | 7.82 GiB | 3.46 GiB | 11.97 GiB | 0.03 GiB | 9±8.3% |
| Llama-3.2-8X3B-MOE-Dark-Champion-Instruct-uncensored-abliterated-18.4BMoE | IQ4_XS | 18.4B | 9.44 GiB | 1.97 GiB | 11.96 GiB | 0.04 GiB | 12±37% |
| Qwen3-VL-32B-Instruct-ultra-uncensored-heretic | I1-IQ1_S | 33.4B | 6.82 GiB | 4.50 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Huihui-Qwen3-VL-32B-Instruct-abliterated | I1-IQ1_S | 33.4B | 6.82 GiB | 4.50 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| ColorGUI-32B | I1-IQ1_S | 33.4B | 6.82 GiB | 4.50 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Qwen3-32B-Uncensored | I1-IQ1_S | 32.8B | 6.82 GiB | 4.50 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Qwen3-32B-abliterated | I1-IQ1_S | 32.8B | 6.82 GiB | 4.50 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| AReaL-boba-2-32B | I1-IQ1_S | 32.8B | 6.82 GiB | 4.50 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Assistant_Pepe_32B | I1-IQ1_S | 32.8B | 6.82 GiB | 4.50 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| NVIDIA-Nemotron-Nano-12B-v2 | Q4_K_M | 12.3B | 6.98 GiB | 4.36 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| gemma-4-A4B-98e-v6-coder-itMoE | IQ4_NL | 20.5B | 10.63 GiB | 0.79 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| gemma-4-A4B-98e-v7-coder-itMoE | IQ4_NL | 20.5B | 10.63 GiB | 0.79 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| gemma-4-A4B-98e-v7-coderx-itMoE | IQ4_NL | 20.5B | 10.63 GiB | 0.79 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| EVA-abliterated-TIES-Qwen2.5-14B | I1-Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Neuron-V1-14B-Instruct | I1-Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Ektome-Qwen2.5-Coder-14B-Instruct-PristinelyUncensored | I1-Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Qwen2.5-14B-Instruct-1M-abliterated | I1-Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| DeepCoder-14B-Preview | Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Deepseeker-Kunou-Qwen2.5-14b | I1-Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| SuperNova-Medius | Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| 14B-Qwen2.5-Kunou-v1 | I1-Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Sugoi-14B-Ultra-HF | I1-Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Qwen2.5-14B-Instruct-abliterated-v2 | Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Qwen2.5-14B-Instruct-Uncensored | Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| Qwen2.5-Coder-14B-Instruct-abliterated | Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| OpenCodeReasoning-Nemotron-14B | Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| DeepSeek-R1-Distill-Qwen-14B-abliterated-v2 | I1-Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
| C1-Tachu | I1-Q4_K_S | 14.8B | 7.98 GiB | 3.38 GiB | 11.96 GiB | 0.04 GiB | 9±8.3% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Questions people ask
- What AI models can a Apple M4 run?
- 1712 of 2118 indexed open-weight models fit a Apple M4 at 65,536 context with q4_0 KV cache, the largest being Qwen3.5-35B-A3B at IQ2_S. That covers text, vision-language, image, video and speech models.
- How much usable memory does a Apple M4 actually have?
- Its nameplate is 16 GB, but about 11.16 GiB is available to a model once driver and compositor overhead is accounted for, and only 12 GB of the pool can be allocated to the GPU at all.
- Is a Apple M4 fast for local AI?
- Its memory bandwidth is 120 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.