Intel · workstation

Arc Pro A40 6GB

Arc Pro A40 6GB has 6 GB of VRAM at 192 GB/s — about 5.58 GiB usable after driver and compositor overhead. 439 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
192 GB/s
96-bit bus
Tensor FP16
dense
TDP
50 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 332audio asr 29audio tts 17vision language 44embedding 14video 3

What fits at 128K context

largest quantization that fits, per model · 439 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
EXAONE-4.0-1.2B-abliteratedIQ4_XS1.5B0.81 GiB3.98 GiB5.57 GiB0.01 GiB21±30%
canary-qwen-2.5bBF162.6B4.73 GiB0.00 GiB5.57 GiB0.01 GiB22±30%
EXAONE-Deep-7.8BQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB22±30%
EXAONE-3.5-7.8B-InstructQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB22±30%
GLM-OCRI1-Q4_11.3B0.54 GiB4.25 GiB5.57 GiB0.01 GiB21±30%
Qwen3-TTS-12Hz-0.6B-BaseQ4_K_M915M4.72 GiB0.00 GiB5.57 GiB0.01 GiB22±30%
VoxCPM2F162.3B4.72 GiB0.00 GiB5.57 GiB0.01 GiB22±30%
Dolphin3.0-Qwen2.5-3bQ6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
Qwen2.5-Coder-3B-Instruct-abliteratedI1-Q6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
GRM-Kerlin-3b-AbliteratedI1-Q6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
Mythos-nanoI1-Q6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
MATE-3BI1-Q6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
Mythos-nano-OBLITERATEDI1-Q6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
Qwen2.5-3B-Instruct-UncensoredI1-Q6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
Nanonets-OCR-sQ6_K3.8B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
Qwen2.5-Coder-3BQ6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
raspberry-3BQ6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
VibeThinker-3B-OBLITERATEDI1-Q6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
VibeThinker-3BQ6_K3.1B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
Fourier-Qwen2.5-VL-3B-0.67I1-Q6_K3.8B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
Qwen2.5-VL-3B-InstructQ6_K3.8B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
jina-embeddings-v4Q6_K3.8B2.36 GiB2.39 GiB5.57 GiB0.01 GiB22±30%
LFM2.5-8B-A1BMoEUD-IQ4_XS8.5B3.97 GiB0.80 GiB5.57 GiB0.01 GiB38±37%
t5-v1_1-xxlQ2_K4.8B4.72 GiB0.00 GiB5.56 GiB0.02 GiB22±30%
gemma-4-E4B-itQ3_K_M8.0B3.78 GiB0.97 GiB5.56 GiB0.02 GiB22±30%
deepseek-coder-5.7bmqa-baseQ5_05.7B3.67 GiB1.06 GiB5.55 GiB0.03 GiB22±30%
gemma-3n-E2B-itQ6_K5.4B3.92 GiB0.82 GiB5.55 GiB0.03 GiB22±30%
Darwin-4B-ChimeraQ5_K_L4.0B2.86 GiB1.87 GiB5.55 GiB0.03 GiB22±30%
G9v3-3BIQ3_XS3.0B1.30 GiB3.45 GiB5.55 GiB0.03 GiB22±30%
Dolphin3.0-Qwen2.5-1.5BF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2.5-1.5B-Instruct-abliteratedF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2.5-1.5B-VibeThinker-heretic-uncensored-abliteratedF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
NEXUS-Coder-OBLITERATEDF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
NEXUS-Coder-AbliteratedF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2.5-1.5B-hereticF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2.5-Math-1.5B-InstructBF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
ShellWhisperer-1.5BF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2.5-1.5BF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2.5-1.5B-InstructF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2.5-Coder-1.5B-InstructF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
PiCo-1BF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2-VL-2B-InstructF162.2B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
FableForge-1.5BF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2-1.5B-InstructBF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2-1.5BBF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Qwen2.5-Coder-1.5BF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
NEXUS-MedicalF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
NEXUS-ScienceF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
NEXUS-LegalF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
NEXUS-CoderF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
NEXUS-FinanceF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
NEXUS-SecurityF161.5B2.88 GiB1.86 GiB5.54 GiB0.04 GiB22±30%
Teuken-7B-instruct-research-v0.4I1-IQ1_M7.5B2.57 GiB2.13 GiB5.53 GiB0.05 GiB22±30%
Vikhr-Gemma-2B-instructQ2_K2.6B1.15 GiB3.57 GiB5.53 GiB0.05 GiB22±30%
Gemmasutra-Mini-2B-v1I1-Q2_K2.6B1.15 GiB3.57 GiB5.53 GiB0.05 GiB22±30%
gemma-2-2b-it-abliteratedQ2_K2.6B1.15 GiB3.57 GiB5.53 GiB0.05 GiB22±30%
gemma-2-2b-itQ2_K2.6B1.15 GiB3.57 GiB5.53 GiB0.05 GiB22±30%
FrickFritz-4BI1-Q4_K_M4.7B2.59 GiB2.13 GiB5.53 GiB0.05 GiB22±30%
qwen3.5-4b-agentic-coder-v4I1-Q4_K_M4.7B2.59 GiB2.13 GiB5.53 GiB0.05 GiB22±30%
Newton-bot-3-VLM-mini-4BQ4_K_M4.7B2.59 GiB2.13 GiB5.53 GiB0.05 GiB22±30%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Arc Pro A40 6GB run?
439 of 2118 indexed open-weight models fit a Arc Pro A40 6GB at 131,072 context with q8_0 KV cache, the largest being EXAONE-4.0-1.2B-abliterated at IQ4_XS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Arc Pro A40 6GB actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Arc Pro A40 6GB fast for local AI?
Its memory bandwidth is 192 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.