Apple · apple

Apple M5

Apple M5 has 32 GB of unified memory at 154 GB/s — about 22.32 GiB usable after driver and compositor overhead. 1998 of 2118 indexed models fit at 4K context with q4_0 KV. Note only 24 GB of its 32 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
32 GB
LPDDR5X-9600
Bandwidth
154 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1716vision language 178audio tts 21image 2video 16audio asr 39embedding 26

What fits at 4K context

largest quantization that fits, per model · 1998 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
HomunculusBF1612.5B23.21 GiB0.18 GiB23.99 GiB0.01 GiB5±8.3%
InternVL3_5-30B-A3BQ6_K30.8B23.38 GiB0.00 GiB23.98 GiB0.02 GiB5±8.3%
OmniAtlas-Qwen3-30B-A3BI1-Q6_K31.7B23.37 GiB0.00 GiB23.96 GiB0.04 GiB5±8.3%
Qwen3-Omni-30B-A3B-CaptionerI1-Q6_K31.7B23.37 GiB0.00 GiB23.96 GiB0.04 GiB5±8.3%
Llama-3_3-Nemotron-Super-49B-v1_5Q3_K_S49.9B20.45 GiB2.81 GiB23.95 GiB0.05 GiB6±8.3%
Valkyrie-49B-v2.1I1-IQ3_S49.9B20.45 GiB2.81 GiB23.95 GiB0.05 GiB6±8.3%
Llama-3_3-Nemotron-Super-49B-v1Q3_K_S49.9B20.45 GiB2.81 GiB23.95 GiB0.05 GiB6±8.3%
Nemotron-Labs-Audex-30B-A3BQ4_K_L32.0B23.35 GiB0.00 GiB23.95 GiB0.05 GiB5±8.3%
Aurora-Code-1MoEI1-Q6_K34.7B23.37 GiB0.02 GiB23.95 GiB0.05 GiB29±37%
Qwen3.5-35B-A3BMoEQ5_K_S36.0B23.33 GiB0.02 GiB23.91 GiB0.09 GiB29±37%
Qwen3.6-35B-A3BMoEQ5_K_S36.0B23.33 GiB0.02 GiB23.91 GiB0.09 GiB29±37%
Gemma-4-Novelist-Eclipse-31BQ5_K_L32.7B22.77 GiB0.51 GiB23.90 GiB0.10 GiB6±8.3%
Gemma-4-31B-StyleTuneQ5_K_L32.7B22.77 GiB0.51 GiB23.90 GiB0.10 GiB6±8.3%
Qwen3-Coder-NextMoEUD-IQ2_M79.7B23.25 GiB0.11 GiB23.89 GiB0.11 GiB32±37%
GLM-4.7-Flash-hereticMoEQ6_K29.9B23.27 GiB0.06 GiB23.89 GiB0.11 GiB23±37%
Darwin-35B-A3B-OpusMoEQ5_K_M36.0B23.30 GiB0.02 GiB23.88 GiB0.12 GiB29±37%
grug-35b-v2MoEQ5_K_M35.1B23.30 GiB0.02 GiB23.88 GiB0.12 GiB29±37%
grug-35bMoEQ5_K_M35.1B23.30 GiB0.02 GiB23.88 GiB0.12 GiB29±37%
WorldSim-Opus-3.6-35B-A3BMoEQ5_K_M35.1B23.30 GiB0.02 GiB23.88 GiB0.12 GiB29±37%
Qwen3.6-35B-A3B-AnkoMoEQ5_K_M35.1B23.30 GiB0.02 GiB23.88 GiB0.12 GiB29±37%
KAT-Coder-V2.5-DevMoEQ5_K_M34.7B23.30 GiB0.02 GiB23.88 GiB0.12 GiB29±37%
Ornith-1.0-35BMoEQ5_K_M34.7B23.30 GiB0.02 GiB23.88 GiB0.12 GiB29±37%
Nex-N2-miniMoEQ5_K_M35.1B23.30 GiB0.02 GiB23.88 GiB0.12 GiB29±37%
Qwen2.5-Coder-32BQ5_132.8B22.95 GiB0.28 GiB23.88 GiB0.12 GiB6±8.3%
Qwen2.5-Coder-32B-InstructQ2_K32.8B22.93 GiB0.28 GiB23.87 GiB0.13 GiB6±8.3%
NVIDIA-Nemotron-Nano-12B-v2BF1612.3B22.94 GiB0.27 GiB23.83 GiB0.17 GiB6±8.3%
Maenad-70BI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
DeepSeek-R1-Distill-Llama-70B-Uncensored-v2-Unbiased-ReasonerI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Rombos-LLM-70b-Llama-3.3I1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
L3.3-Electra-R1-70bI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Latxa-Llama-3.1-70B-Instruct-v2I1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Llama-3.3_70_b_uncensored_continuedI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Llama-3.3-70B-Instruct-abliteratedI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
grok-oss-Revenant-70BI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Llama-3.1-Nemotron-70B-Instruct-HFI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
L3.3-70B-Euryale-v2.3I1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Hermes-4-70B-hereticI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Llama-3.1-70BQ2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Golem-70B-v1bI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
DeepSeek-R1-Distill-Llama-70B-abliteratedI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
DeepSeek-R1-Distill-Llama-70B-hereticI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Legion-V2.1-LLaMa-70BI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Assistant_Pepe_70BI1-Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Athene-70BQ2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Hermes-3-Llama-3.1-70BQ2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
L3.3-70B-Magnum-DiamondQ2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Meta-Llama-3-70B-Instruct-abliterated-v3.5Q2_K_S70.6B22.79 GiB0.35 GiB23.82 GiB0.18 GiB6±8.3%
Laguna-S-2.1MoEIQ1_S118B23.15 GiB0.09 GiB23.81 GiB0.19 GiB27±37%
Qwen-AgentWorld-35B-A3BMoEUD-Q5_K_S34.7B23.23 GiB0.02 GiB23.81 GiB0.19 GiB30±37%
WizardLM-Uncensored-SuperCOT-StoryTelling-30bQ5_K_M32.5B21.46 GiB1.71 GiB23.80 GiB0.20 GiB6±8.3%
Wizard-Vicuna-30B-UncensoredI1-Q5_K_M32.5B21.46 GiB1.71 GiB23.80 GiB0.20 GiB6±8.3%
archangel_sft-kto_llama30bI1-Q5_K_M32.5B21.46 GiB1.71 GiB23.80 GiB0.20 GiB6±8.3%
ALIA-40b-fc-2606I1-Q4_K_M40.4B22.90 GiB0.21 GiB23.78 GiB0.22 GiB6±8.3%
ALIA-40b-instruct-2606I1-Q4_K_M40.4B22.90 GiB0.21 GiB23.78 GiB0.22 GiB6±8.3%
Qwen3-Coder-Next-REAMMoEI1-IQ3_XS60.3B23.19 GiB0.03 GiB23.76 GiB0.24 GiB31±37%
Le-Chaton-Slim-23BMoEQ8_023.3B23.07 GiB0.11 GiB23.75 GiB0.25 GiB12±37%
Nemotron-Cascade-2-30B-A3BMoEQ4_K_L31.6B23.15 GiB0.06 GiB23.74 GiB0.26 GiB26±37%
Llama3.2-30B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEI1-Q6_K30.0B22.97 GiB0.21 GiB23.74 GiB0.26 GiB16±37%
Qwen3.5-122B-A10B-hereticMoEI1-IQ1_S123B23.13 GiB0.03 GiB23.74 GiB0.26 GiB30±37%
DeepSeek-R1-Distill-Llama-70BUD-IQ2_M70.6B22.70 GiB0.35 GiB23.72 GiB0.28 GiB6±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Apple M5 run?
1998 of 2118 indexed open-weight models fit a Apple M5 at 4,096 context with q4_0 KV cache, the largest being Homunculus at BF16. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M5 actually have?
Its nameplate is 32 GB, but about 22.32 GiB is available to a model once driver and compositor overhead is accounted for, and only 24 GB of the pool can be allocated to the GPU at all.
Is a Apple M5 fast for local AI?
Its memory bandwidth is 154 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.