Apple · apple

Apple M1

Apple M1 has 16 GB of unified memory at 68 GB/s — about 11.16 GiB usable after driver and compositor overhead. 1811 of 2118 indexed models fit at 8K context with q8_0 KV. Note only 12 GB of its 16 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
LPDDR4X-4266
Bandwidth
68 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
video 15text 1557vision language 151audio asr 39audio tts 21image 2embedding 26

What fits at 8K context

largest quantization that fits, per model · 1811 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Wan2.1-VACE-14BQ5_K_S17.3B11.41 GiB0.00 GiB12.00 GiB0.00 GiB5±8.3%
Goetia-26B-A4B-v1.4MoEI1-IQ3_XS26.0B11.13 GiB0.32 GiB11.99 GiB0.01 GiB5±8.3%
G4-Moonlight-Dusk-26B-A4B-hereticMoEI1-IQ3_XS26.5B11.13 GiB0.32 GiB11.99 GiB0.01 GiB5±8.3%
Pantheon-Reasoning-26B-A4B-1.1-hereticMoEI1-IQ3_XS26.5B11.13 GiB0.32 GiB11.99 GiB0.01 GiB5±8.3%
G4-Moonlight-Dusk-26B-A4BMoEI1-IQ3_XS26.5B11.13 GiB0.32 GiB11.99 GiB0.01 GiB5±8.3%
Chimera-X-26B-A4BMoEI1-IQ3_XS26.5B11.13 GiB0.32 GiB11.99 GiB0.01 GiB5±8.3%
Pantheon-Reasoning-26B-A4B-1.1MoEI1-IQ3_XS26.5B11.13 GiB0.32 GiB11.99 GiB0.01 GiB5±8.3%
Gemma-4-26B-A4B-StyleTune-V2MoEI1-IQ3_XS26.5B11.13 GiB0.32 GiB11.99 GiB0.01 GiB5±8.3%
Gemma-4-26B-A4B-StyleTuneMoEI1-IQ3_XS26.5B11.13 GiB0.32 GiB11.99 GiB0.01 GiB5±8.3%
gemma-4-26b-a4b-heretic-styletune-v2-headMoEI1-IQ3_XS25.8B11.13 GiB0.32 GiB11.99 GiB0.01 GiB5±8.3%
internlm2-math-plus-20bI1-Q4_019.9B10.58 GiB0.80 GiB11.99 GiB0.01 GiB5±8.3%
Kimi-Linear-48B-A3B-InstructMoEIQ2_XXS49.1B11.30 GiB0.13 GiB11.98 GiB0.02 GiB5±8.3%
Salience-1.5-FlashMoEI1-IQ3_XXS31.1B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Huihui-Qwen3-VL-30B-A3B-Instruct-abliteratedMoEI1-IQ3_XXS31.1B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Qwen3-30B-A3B-Gemini-Pro-High-Reasoning-2507-ABLITERATED-UNCENSOREDMoEI1-IQ3_XXS30.5B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
MiroThinker-v1.0-30BMoEI1-IQ3_XXS30.5B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Qwen3-30B-A3B-YOYO-V5MoEI1-IQ3_XXS30.5B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Qwen3-30B-A3B-Thinking-2507-Claude-4.5-Sonnet-High-Reasoning-DistillMoEI1-IQ3_XXS30.5B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Huihui-Qwen3-30B-A3B-Thinking-2507-abliteratedMoEI1-IQ3_XXS30.5B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Huihui-Qwen3-30B-A3B-Instruct-2507-abliteratedMoEI1-IQ3_XXS30.5B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Qwen3-30B-A3B-abliterated-eroticMoEI1-IQ3_XXS30.5B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Huihui-Qwen3-Coder-30B-A3B-Instruct-abliteratedMoEI1-IQ3_XXS30.5B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Qwen3-Coder-30B-A3B-Instruct-RTPurboMoEI1-IQ3_XXS30.5B11.04 GiB0.40 GiB11.98 GiB0.02 GiB18±37%
Qwen3.8-27BUD-IQ3_XXS27.8B11.10 GiB0.27 GiB11.97 GiB0.03 GiB5±8.3%
InternVL3_5-30B-A3BIQ3_XXS30.8B11.38 GiB0.00 GiB11.97 GiB0.03 GiB5±8.3%
gpt-oss-20b-hereticMoEIQ3_XS20.9B11.32 GiB0.11 GiB11.97 GiB0.03 GiB14±37%
OpenAI-gpt-oss-20B-Claude-4.5-Opus-Heretic-UncensoredMoEI1-Q4_020.9B11.31 GiB0.11 GiB11.96 GiB0.04 GiB14±37%
gpt-oss-20b-uncensoredMoEI1-Q4_020.9B11.31 GiB0.11 GiB11.96 GiB0.04 GiB14±37%
gpt-oss-safeguard-20bMoEI1-Q4_021.5B11.31 GiB0.11 GiB11.96 GiB0.04 GiB14±37%
Huihui-gpt-oss-20b-BF16-abliterated-v2MoEI1-Q4_020.9B11.31 GiB0.11 GiB11.96 GiB0.04 GiB14±37%
metatune-gpt20b-R1.09MoEI1-Q4_021.5B11.31 GiB0.11 GiB11.96 GiB0.04 GiB14±37%
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresyMoEI1-Q3_K_L23.0B11.18 GiB0.22 GiB11.96 GiB0.04 GiB17±37%
Carnice-Qwen3.6-MoE-35B-A3BMoEI1-Q2_K_S36.0B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Qwen35B-Agent-R2-AbliteratedMoEI1-Q2_K_S34.7B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoEI1-Q2_K_S36.0B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Darwin-35B-A3B-OpusMoEI1-Q2_K_S36.0B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Qwen35B-Agent-R2MoEI1-Q2_K_S34.7B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Carnice-MoE-35B-A3BMoEI1-Q2_K_S36.0B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
spoomplesmaxx-flash-35B-A3MoEI1-Q2_K_S35.1B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Huihui-Qwen3.6-35B-A3B-Claude-4.7-Opus-abliteratedMoEI1-Q2_K_S36.0B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Qwen3.6-35B-A3B-Uncensored-AggressiveMoEI1-Q2_K_S35.1B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
WorldSim-Opus-3.6-35B-A3BMoEI1-Q2_K_S35.1B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Qwen3.6-35B-A3B-abliterated-MAXMoEI1-Q2_K_S35.1B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Huihui-Qwen3.6-35B-A3B-abliteratedMoEI1-Q2_K_S36.0B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Qwopus3.6-35B-A3B-v1MoEI1-Q2_K_S36.0B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Qwen3.6-35B-A3B-StyleTuneMoEI1-Q2_K_S35.1B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Qwen3.6-35B-A3B-abliteratedMoEI1-Q2_K_S35.1B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
0GM-1.0-35B-A3B-0427MoEI1-Q2_K_S36.0B11.32 GiB0.08 GiB11.95 GiB0.05 GiB26±37%
Mistral-MOE-4X7B-Dark-MultiVerse-Uncensored-Enhanced32-24BMoEQ3_K_M24.2B10.83 GiB0.53 GiB11.95 GiB0.05 GiB3±37%
Trinity-MiniMoEIQ3_M26.1B11.27 GiB0.13 GiB11.94 GiB0.06 GiB20±37%
Gemma-4-31B-Isometry-RPI1-IQ2_S32.7B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
Gemma-4-Dark-Gemistry-31BI1-IQ2_S32.7B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
Prosopon-31BI1-IQ2_S32.7B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
Gemma-4-Novelist-Eclipse-31BI1-IQ2_S32.7B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
Giftige-Blume-31B-v1-StyleSwapI1-IQ2_S32.7B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
G4-MeroMero-31B-StyleSwapI1-IQ2_S32.7B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
Gemma-4-31B-StyleTune-heretic-araI1-IQ2_S32.7B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
Pantheon-Reasoning-31B-1.1I1-IQ2_S32.7B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
Gemma-4-31B-StyleTuneI1-IQ2_S32.7B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
Barcenas-StyleTune-31B-FableI1-IQ2_S32.1B10.02 GiB1.29 GiB11.94 GiB0.06 GiB5±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M1 run?
1811 of 2118 indexed open-weight models fit a Apple M1 at 8,192 context with q8_0 KV cache, the largest being Wan2.1-VACE-14B at Q5_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M1 actually have?
Its nameplate is 16 GB, but about 11.16 GiB is available to a model once driver and compositor overhead is accounted for, and only 12 GB of the pool can be allocated to the GPU at all.
Is a Apple M1 fast for local AI?
Its memory bandwidth is 68 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.