Apple · apple

Apple M4

Apple M4 has 16 GB of unified memory at 120 GB/s — about 11.16 GiB usable after driver and compositor overhead. 1497 of 2118 indexed models fit at 128K context with q4_0 KV. Note only 12 GB of its 16 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
LPDDR5X-7500
Bandwidth
120 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1262video 15vision language 134audio tts 21embedding 26audio asr 38image 1

What fits at 128K context

largest quantization that fits, per model · 1497 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
DeepSeek-V2-Lite-Chat-UncensoredMoEQ5_K_S15.7B10.37 GiB1.07 GiB12.00 GiB0.00 GiB20±37%
Qwen3.6-27B-A3B-CoderMoEI1-Q3_K_S26.7B10.74 GiB0.70 GiB12.00 GiB0.00 GiB26±37%
INTELLECT-1-InstructI1-Q4_010.2B5.50 GiB5.91 GiB12.00 GiB0.00 GiB9±8.3%
Wan2.1-VACE-14BQ5_K_S17.3B11.41 GiB0.00 GiB12.00 GiB0.00 GiB9±8.3%
Fimbulvetr-11B-v2I1-IQ3_M10.7B4.66 GiB6.75 GiB11.99 GiB0.01 GiB9±8.3%
gemma-4-26B-A4B-itMoEIQ2_M26.5B9.97 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
dolphincoder-starcoder2-15bKV unresolvedI1-Q4_K_S16.0B8.53 GiB2.81 GiB11.99 GiB0.01 GiB9±8.3%
Goetia-26B-A4B-v1.4MoEI1-IQ2_M26.0B9.96 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
G4-Moonlight-Dusk-26B-A4B-hereticMoEI1-IQ2_M26.5B9.96 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEI1-IQ2_M26.5B9.96 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
G4-Moonlight-Dusk-26B-A4BMoEI1-IQ2_M26.5B9.96 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
Chimera-X-26B-A4BMoEI1-IQ2_M26.5B9.96 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
Pantheon-Reasoning-26B-A4B-1.1MoEI1-IQ2_M26.5B9.96 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
Gemma-4-26B-A4B-StyleTune-V2MoEI1-IQ2_M26.5B9.96 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
Gemma-4-26B-A4B-StyleTuneMoEI1-IQ2_M26.5B9.96 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEI1-IQ2_M25.8B9.96 GiB1.49 GiB11.99 GiB0.01 GiB9±8.3%
Qwen3-VL-30B-A3B-ThinkingMoEIQ2_XS31.1B8.07 GiB3.38 GiB11.99 GiB0.01 GiB13±37%
MiroThinker-v1.0-30BMoEIQ2_XS30.5B8.07 GiB3.38 GiB11.99 GiB0.01 GiB13±37%
Qwen3-30B-A3B-Instruct-2507MoEIQ2_XS30.5B8.07 GiB3.38 GiB11.99 GiB0.01 GiB13±37%
Qwen3-30B-A3B-Thinking-2507MoEIQ2_XS30.5B8.07 GiB3.38 GiB11.99 GiB0.01 GiB13±37%
Tongyi-DeepResearch-30B-A3BMoEIQ2_XS30.5B8.07 GiB3.38 GiB11.98 GiB0.02 GiB13±37%
North-Mini-Code-1.0MoEQ2_K_L30.5B10.45 GiB1.00 GiB11.98 GiB0.02 GiB23±37%
Mistral-7B-v0.1KV unresolvedQ5_07.2B6.89 GiB4.50 GiB11.98 GiB0.02 GiB9±8.3%
gemma-3-12b-it-vl-Gemini-3-Pro-Preview-Heretic-Uncensored-ThinkingI1-Q6_K12.2B9.00 GiB2.38 GiB11.97 GiB0.03 GiB9±8.3%
gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-ThinkingI1-Q6_K12.2B9.00 GiB2.38 GiB11.97 GiB0.03 GiB9±8.3%
gemma-3-12b-it-ultra-uncensored-hereticQ6_K12.2B9.00 GiB2.38 GiB11.97 GiB0.03 GiB9±8.3%
gemma-3-12b-it-vl-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-Q6_K12.2B9.00 GiB2.38 GiB11.97 GiB0.03 GiB9±8.3%
Floppa-12B-Gemma3-UncensoredI1-Q6_K12.2B9.00 GiB2.38 GiB11.97 GiB0.03 GiB9±8.3%
gemma-3-12b-it-hereticI1-Q6_K12.2B9.00 GiB2.38 GiB11.97 GiB0.03 GiB9±8.3%
gemma-3-12b-it-abliteratedQ6_K12.2B9.00 GiB2.38 GiB11.97 GiB0.03 GiB9±8.3%
gemma-3-12b-it-abliterated-v2Q6_K11.8B9.00 GiB2.38 GiB11.97 GiB0.03 GiB9±8.3%
gemma-3-12b-itQ6_K12.2B9.00 GiB2.38 GiB11.97 GiB0.03 GiB9±8.3%
InternVL3_5-30B-A3BIQ3_XXS30.8B11.38 GiB0.00 GiB11.97 GiB0.03 GiB9±8.3%
Qwen3.5-21B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingI1-Q3_K_M21.3B9.67 GiB1.69 GiB11.97 GiB0.03 GiB9±8.3%
Qwen3.6-21B-IQ-Ultra-Heretic-Uncensored-ThinkingI1-Q3_K_M21.3B9.67 GiB1.69 GiB11.97 GiB0.03 GiB9±8.3%
Qwen-AgentWorld-35B-A3BMoEUD-IQ2_XXS34.7B10.71 GiB0.70 GiB11.96 GiB0.04 GiB29±37%
Ornith-1.0-35BMoEUD-IQ2_XXS34.7B10.71 GiB0.70 GiB11.96 GiB0.04 GiB29±37%
Salience-1.5-ProMoEIQ2_S36.0B10.70 GiB0.70 GiB11.96 GiB0.04 GiB29±37%
Qwable-v1MoEIQ2_S36.0B10.70 GiB0.70 GiB11.96 GiB0.04 GiB29±37%
T-SearchMoEIQ2_S36.0B10.70 GiB0.70 GiB11.96 GiB0.04 GiB29±37%
DeepSeek-R1-Distill-Llama-8B-AbliteratedI1-IQ3_S8.0B6.86 GiB4.50 GiB11.95 GiB0.05 GiB9±8.3%
GLM-4.7-Flash-DerestrictedMoEI1-Q2_K_S31.2B9.53 GiB1.86 GiB11.95 GiB0.05 GiB18±37%
Huihui-GLM-4.7-Flash-abliteratedMoEI1-Q2_K_S31.2B9.53 GiB1.86 GiB11.95 GiB0.05 GiB18±37%
internlm2-math-plus-20bI1-IQ1_M19.9B4.58 GiB6.75 GiB11.94 GiB0.06 GiB9±8.3%
gemma-4-19B-A4B-it-INSTRUCT-Heretic-UncensoredMoEI1-Q4_019.0B9.92 GiB1.49 GiB11.94 GiB0.06 GiB9±8.3%
gemma-4-19B-A4B-it-The-DECKARD-Heretic-Uncensored-ThinkingMoEI1-Q4_019.0B9.92 GiB1.49 GiB11.94 GiB0.06 GiB9±8.3%
gemma-4-19b-a4b-it-REAP-hereticMoEI1-Q4_019.0B9.92 GiB1.49 GiB11.94 GiB0.06 GiB9±8.3%
Gemma-4-19BMoEI1-Q4_019.0B9.92 GiB1.49 GiB11.94 GiB0.06 GiB9±8.3%
grug-27bIQ2_XS27.4B9.08 GiB2.25 GiB11.94 GiB0.06 GiB9±8.3%
Carnice-V2-27bIQ2_XS27.4B9.08 GiB2.25 GiB11.94 GiB0.06 GiB9±8.3%
Trinity-MiniMoEQ3_K_S26.1B10.80 GiB0.60 GiB11.94 GiB0.06 GiB26±37%
Falcon3-7B-InstructQ8_07.5B7.38 GiB3.94 GiB11.94 GiB0.06 GiB9±8.3%
Wan2.1-T2V-1.3BQ8_01.4B11.37 GiB0.00 GiB11.93 GiB0.07 GiB9±8.3%
Moonlight-16B-A3B-InstructMoEQ5_016.0B10.30 GiB1.07 GiB11.93 GiB0.07 GiB20±37%
granite-3.3-8b-instructQ5_18.2B5.72 GiB5.63 GiB11.93 GiB0.07 GiB9±8.3%
granite-3.2-8b-instructQ5_18.2B5.72 GiB5.63 GiB11.93 GiB0.07 GiB9±8.3%
Qwen3-Coder-REAP-25B-A3BMoEQ2_K_S24.9B8.01 GiB3.38 GiB11.92 GiB0.08 GiB12±37%
Qwen3-VL-8B-Instruct-HereticI1-IQ3_XXS8.8B6.28 GiB5.06 GiB11.92 GiB0.08 GiB9±8.3%
Goetia-26B-A4B-v1.3-Absolute-Heretic-ARAMoEI1-Q2_K_S25.8B9.89 GiB1.49 GiB11.92 GiB0.08 GiB9±8.3%
Frank-26B-A4BMoEI1-Q2_K_S26.5B9.89 GiB1.49 GiB11.92 GiB0.08 GiB9±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M4 run?
1497 of 2118 indexed open-weight models fit a Apple M4 at 131,072 context with q4_0 KV cache, the largest being DeepSeek-V2-Lite-Chat-Uncensored at Q5_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M4 actually have?
Its nameplate is 16 GB, but about 11.16 GiB is available to a model once driver and compositor overhead is accounted for, and only 12 GB of the pool can be allocated to the GPU at all.
Is a Apple M4 fast for local AI?
Its memory bandwidth is 120 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.