Apple · apple

Apple M2

Apple M2 has 8 GB of unified memory at 102 GB/s — about 5.58 GiB usable after driver and compositor overhead. 1195 of 2118 indexed models fit at 16K context with q8_0 KV. Note only 6 GB of its 8 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
LPDDR5-6400
Bandwidth
102 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 95text 1012embedding 26audio asr 38audio tts 19video 5

What fits at 16K context

largest quantization that fits, per model · 1195 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
LocateAnything-3BQ6_K3.8B5.13 GiB0.30 GiB5.99 GiB0.01 GiB15±8.3%
salamandra-7b-instruct-2606I1-Q4_K_S7.8B4.35 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
saiga_llama3_8bQ4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
MiniCPM-Llama3-V-2_5Q4_08.5B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Llama3-ChatQA-1.5-8BQ4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Mistral-7B-v0.1KV unresolvedQ3_K_S7.2B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Llama-3-Groq-8B-Tool-UseQ4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Dolphin3.0-Llama3.1-8B-abliteratedQ4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Llama-3.1-8BQ4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Llama-3.1-8B-Lexi-Uncensored-V2Q4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Llama-3.1-Swallow-8B-Instruct-v0.5Q4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Turkish-Llama-8b-Instruct-v0.1Q4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Infinity-Instruct-7M-Gen-Llama3_1-8BQ4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Meta-Llama-3-8B-InstructQ4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
openchat-3.6-8b-20240522Q4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
L3-8B-Stheno-v3.3-32KQ4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
NeuralDaredevil-8B-abliteratedQ4_08.0B4.34 GiB1.06 GiB5.99 GiB0.01 GiB15±8.3%
Gemma-The-Writer-N-Restless-Quill-10B-UncensoredI1-IQ2_M10.0B3.44 GiB1.96 GiB5.99 GiB0.01 GiB15±8.3%
gemma-2-2b-it-abliteratedF162.6B4.88 GiB0.55 GiB5.99 GiB0.01 GiB15±8.3%
Vikhr-Gemma-2B-instructBF162.6B4.88 GiB0.55 GiB5.99 GiB0.01 GiB15±8.3%
octo-netQ4_13.8B2.24 GiB3.19 GiB5.99 GiB0.01 GiB15±8.3%
Qwen3-VL-8B-Instruct-HereticI1-IQ1_M8.8B4.20 GiB1.20 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-uncensoredI1-Q5_K_S7.9B5.26 GiB0.15 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-qat-q4_0-unquantized-hereticI1-Q5_K_S7.9B5.26 GiB0.15 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-qat-heretic_decensoredI1-Q5_K_S7.9B5.26 GiB0.15 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-QAT-SOMPOA-heresyI1-Q5_K_S7.9B5.26 GiB0.15 GiB5.98 GiB0.02 GiB15±8.3%
gemma4-e4b-mahou-nsfwI1-Q5_K_S7.9B5.26 GiB0.15 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-mentalchat16kI1-Q5_K_S7.9B5.26 GiB0.15 GiB5.98 GiB0.02 GiB15±8.3%
gemma4-E4B-it-abliteratedI1-Q5_K_S7.9B5.26 GiB0.15 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-OBLITERATEDI1-Q5_K_S8.0B5.26 GiB0.15 GiB5.98 GiB0.02 GiB15±8.3%
Grug-12BIQ2_M12.0B4.60 GiB0.78 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-12B-it-Esper4IQ2_M12.0B4.60 GiB0.78 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-12B-itIQ2_M12.0B4.60 GiB0.78 GiB5.98 GiB0.02 GiB15±8.3%
Aya-Medikal-V2IQ4_XS8.0B4.32 GiB1.06 GiB5.97 GiB0.03 GiB15±8.3%
stable-code-3bQ8_02.8B2.77 GiB2.66 GiB5.97 GiB0.03 GiB15±8.3%
rocket-3BQ8_02.8B2.77 GiB2.66 GiB5.97 GiB0.03 GiB15±8.3%
NuExtract-1.5Q4_K_M3.8B2.23 GiB3.19 GiB5.97 GiB0.03 GiB15±8.3%
Phi-3.5-mini-instructQ4_K_M3.8B2.23 GiB3.19 GiB5.97 GiB0.03 GiB15±8.3%
Phi-3.5-mini-instruct_UncensoredQ4_K_M3.8B2.23 GiB3.19 GiB5.97 GiB0.03 GiB15±8.3%
Phi-3-mini-128k-instructQ4_K_M3.8B2.23 GiB3.19 GiB5.97 GiB0.03 GiB15±8.3%
Phi-3-mini-4k-instructQ4_K_M3.8B2.23 GiB3.19 GiB5.97 GiB0.03 GiB15±8.3%
Chocolatine-3B-Instruct-DPO-RevisedQ4_K_M3.8B2.23 GiB3.19 GiB5.97 GiB0.03 GiB15±8.3%
phi-2Q8_02.8B2.75 GiB2.66 GiB5.97 GiB0.03 GiB15±8.3%
Ministral-3-8B-Instruct-2512-BF16-abliteratedI1-Q3_K_L8.9B4.25 GiB1.13 GiB5.97 GiB0.03 GiB15±8.3%
Amaretto-8BI1-Q3_K_L8.9B4.25 GiB1.13 GiB5.97 GiB0.03 GiB15±8.3%
zeta-2.1IQ4_XS8.3B4.31 GiB1.06 GiB5.97 GiB0.03 GiB15±8.3%
Ornith-1.0-9BIQ4_NL9.2B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
Qwen3.6-35B-A3B-REAM-160-ru-agentMoEIQ1_S23.6B5.24 GiB0.17 GiB5.96 GiB0.04 GiB49±37%
Falcon3-10B-InstructQ2_K_L10.3B4.02 GiB1.33 GiB5.96 GiB0.04 GiB15±8.3%
internlm3-8b-instructQ4_K_M8.8B4.99 GiB0.40 GiB5.96 GiB0.04 GiB15±8.3%
Qwen3.5-9B-CoderI1-Q4_K_S9.7B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
Qwythos-9B-Claude-Mythos-5-1M-MTPI1-Q4_K_S9.7B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliteratedI1-Q4_K_S9.7B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
Qwen3.5-9B-Fable-5-v1I1-Q4_K_S9.7B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
Qwythos-9B-v2I1-Q4_K_S9.7B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
PINQWEN-3.5-9B-1M-BF16I1-Q4_K_S9.7B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
Openprose-2-FlashI1-Q4_K_S9.7B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
Qwen3.5-9B-Nikusui-v1I1-Q4_K_S9.7B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
Ornstein-3.5-9B-V1.5I1-Q4_K_S9.7B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
Ornith-1.0-9B-heretic-MTPI1-Q4_K_S9.4B5.11 GiB0.27 GiB5.96 GiB0.04 GiB15±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing147.27 tok/s115.58180.497
Text generation12.18 tok/s7.6716.967
Benchmarked· n=7

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.

Questions people ask

What AI models can a Apple M2 run?
1195 of 2118 indexed open-weight models fit a Apple M2 at 16,384 context with q8_0 KV cache, the largest being LocateAnything-3B at Q6_K. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M2 actually have?
Its nameplate is 8 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for, and only 6 GB of the pool can be allocated to the GPU at all.
Is a Apple M2 fast for local AI?
Its memory bandwidth is 102 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.