Apple · apple

Apple M2

Apple M2 has 8 GB of unified memory at 102 GB/s — about 5.58 GiB usable after driver and compositor overhead. 1260 of 2118 indexed models fit at 16K context with q4_0 KV. Note only 6 GB of its 8 GB is allocatable to the GPU.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
LPDDR5-6400
Bandwidth
102 GB/s
128-bit bus
Tensor FP16
dense
TDP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1076vision language 95embedding 26image 1audio asr 38audio tts 19video 5

What fits at 16K context

largest quantization that fits, per model · 1260 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Teuken-7B-instruct-research-v0.4I1-Q5_K_M7.5B5.27 GiB0.14 GiB6.00 GiB0.00 GiB15±8.3%
Mistral-7B-v0.1KV unresolvedQ3_K_M7.2B4.84 GiB0.56 GiB5.99 GiB0.01 GiB15±8.3%
dolphin-2.9.2-Phi-3-MediumKV unresolvedQ2_K_S14.0B4.50 GiB0.88 GiB5.99 GiB0.01 GiB15±8.3%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopIQ3_S12.1B5.27 GiB0.11 GiB5.99 GiB0.01 GiB15±8.3%
Qwen3.5-9BIQ4_NL9.7B5.26 GiB0.14 GiB5.98 GiB0.02 GiB15±8.3%
Gemma-The-Writer-N-Restless-Quill-10B-UncensoredI1-IQ3_S10.0B4.36 GiB1.04 GiB5.98 GiB0.02 GiB15±8.3%
Anubis-Mini-8B-v1Q4_18.0B4.83 GiB0.56 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-uncensoredI1-Q5_K_M7.9B5.33 GiB0.08 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-qat-q4_0-unquantized-hereticI1-Q5_K_M7.9B5.33 GiB0.08 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-qat-heretic_decensoredI1-Q5_K_M7.9B5.33 GiB0.08 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-QAT-SOMPOA-heresyI1-Q5_K_M7.9B5.33 GiB0.08 GiB5.98 GiB0.02 GiB15±8.3%
gemma4-e4b-mahou-nsfwI1-Q5_K_M7.9B5.33 GiB0.08 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-mentalchat16kI1-Q5_K_M7.9B5.33 GiB0.08 GiB5.98 GiB0.02 GiB15±8.3%
gemma4-E4B-it-abliteratedI1-Q5_K_M7.9B5.33 GiB0.08 GiB5.98 GiB0.02 GiB15±8.3%
gemma-4-E4B-it-OBLITERATEDI1-Q5_K_M8.0B5.33 GiB0.08 GiB5.98 GiB0.02 GiB15±8.3%
Huihui-gemma-3n-E4B-it-abliteratedQ6_K7.8B5.31 GiB0.08 GiB5.97 GiB0.03 GiB15±8.3%
legitus-instruct-v1I1-Q4_18.1B4.79 GiB0.56 GiB5.97 GiB0.03 GiB15±8.3%
Apertus-8B-Instruct-2509I1-Q4_18.1B4.79 GiB0.56 GiB5.97 GiB0.03 GiB15±8.3%
Smilodon-9B-v1I1-Q3_K_M10.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
bella-bartender-v2I1-Q3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-Q3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-Q3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
Gemma-2-9B-It-SPPO-Iter3I1-Q3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
Gemma-SEA-LION-v3-9B-ITI1-Q3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-Q3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
G2-Darkest-Writer-9B-v1I1-Q3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
Tiger-Gemma-9B-v3I1-Q3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
gemma-2-9b-it-abliteratedQ3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
gemma-2-9b-itQ3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
Tiger-Gemma-9B-v1Q3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
magnum-v4-9bQ3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
gemma-2-9bQ3_K_M9.2B4.43 GiB0.95 GiB5.97 GiB0.03 GiB15±8.3%
Vero-Qwen35-9B-BaseI1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Vero-Qwen35-9BI1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-ThinkingI1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Morphos-9BI1-Q4_K_M9.0B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwable-9B-Claude-Fable-5-hereticI1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Holo-3.1-9BI1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwable-9B-Claude-Fable-5I1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwen3.5-9B-imabari-v2I1-Q4_K_M9.7B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwen3.5-9B-Fable-5-Quad-StockQ4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwen3.5-9B-abliterated-v2-MAXI1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
OmniCoder-9B-Claude-Opus-High-Reasoning-DistillI1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwable-9B-Claude-Fable-5-StraTAI1-Q4_K_M9.0B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwable-9B-Claude-Fable-5-OBLITERATEDI1-Q4_K_M9.0B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwen3.5-9B-RpRMax-v1I1-Q4_K_M9.7B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
AdQWENistrator-9BI1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
cajal-9b-v2-fullI1-Q4_K_M9.0B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwen3.5-9B-ultra-uncensored-hereticQ4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Holo-3.1-9B-CoderI1-Q4_K_M9.0B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
PlutoI1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Holo-3.1-9B-abliterated-rdoI1-Q4_K_M9.0B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwen3.5-9B-Uncensored-cyber-v3Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwen3.5-9B-BaseI1-Q4_K_M9.7B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
qwen3.5-9b-nsfw-captioning-v5I1-Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Miss_MARTHA-9B-Qwen3.5-OmniI1-Q4_K_M9.0B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
MaralGPT-Mythos-9B-2606Q4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
QwenPaw-Flash-9B-hereticQ4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Qwen3.5-9B-hereticQ4_K_M9.4B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
Huihui-Qwen3.5-9B-abliteratedQ4_K_M9.7B5.24 GiB0.14 GiB5.97 GiB0.03 GiB15±8.3%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Prompt processing147.27 tok/s115.58180.497
Text generation12.18 tok/s7.6716.967
Benchmarked· n=7

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from llama.cpp-discussion-4167.

Questions people ask

What AI models can a Apple M2 run?
1260 of 2118 indexed open-weight models fit a Apple M2 at 16,384 context with q4_0 KV cache, the largest being Teuken-7B-instruct-research-v0.4 at I1-Q5_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Apple M2 actually have?
Its nameplate is 8 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for, and only 6 GB of the pool can be allocated to the GPU at all.
Is a Apple M2 fast for local AI?
Its memory bandwidth is 102 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.