Intel · consumer

Arc A310 4GB

Arc A310 4GB has 4 GB of VRAM at 124 GB/s — about 3.72 GiB usable after driver and compositor overhead. 802 of 2118 indexed models fit at 8K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
4 GB
GDDR6
Bandwidth
124 GB/s
64-bit bus
Tensor FP16
dense
TDP
75 W
$110 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 59text 660embedding 25audio asr 37video 2audio tts 19

What fits at 8K context

largest quantization that fits, per model · 802 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
ToriiGate-0.5IQ4_NL5.2B2.77 GiB0.13 GiB3.72 GiB0.00 GiB23±30%
NuExtract-1.5Q2_K3.8B1.32 GiB1.59 GiB3.72 GiB0.00 GiB23±30%
Phi-3.5-mini-instructQ2_K3.8B1.32 GiB1.59 GiB3.72 GiB0.00 GiB23±30%
Phi-3.5-mini-instruct_UncensoredQ2_K3.8B1.32 GiB1.59 GiB3.72 GiB0.00 GiB23±30%
Phi-3-mini-128k-instructQ2_K3.8B1.32 GiB1.59 GiB3.72 GiB0.00 GiB23±30%
Phi-3-mini-4k-instructQ2_K3.8B1.32 GiB1.59 GiB3.72 GiB0.00 GiB23±30%
octo-netQ2_K3.8B1.32 GiB1.59 GiB3.72 GiB0.00 GiB23±30%
CycleGRPO-4BI1-IQ4_XS4.8B2.31 GiB0.60 GiB3.72 GiB0.00 GiB23±30%
stable-code-3bI1-Q4_K_M2.8B1.59 GiB1.33 GiB3.72 GiB0.00 GiB23±30%
rocket-3BQ4_K_M2.8B1.59 GiB1.33 GiB3.72 GiB0.00 GiB23±30%
Jan-v3-4B-base-instructIQ4_XS4.4B2.31 GiB0.60 GiB3.72 GiB0.00 GiB23±30%
Jan-code-4bIQ4_XS4.4B2.31 GiB0.60 GiB3.72 GiB0.00 GiB23±30%
Miril-Drone-2B-1IQ3_XS5.1B2.88 GiB0.04 GiB3.72 GiB0.00 GiB23±30%
LFM2-8B-A1BMoEQ2_K8.3B2.87 GiB0.05 GiB3.71 GiB0.01 GiB62±37%
Nanbeige4.1-3BQ5_K_M3.9B2.63 GiB0.27 GiB3.71 GiB0.01 GiB23±30%
Yi-6B-ChatI1-IQ3_M6.1B2.62 GiB0.27 GiB3.71 GiB0.01 GiB23±30%
rnj-1-instructUD-IQ2_XXS8.3B2.33 GiB0.53 GiB3.71 GiB0.01 GiB23±30%
Luna-7B-A4BMoEI1-Q2_K_S6.7B2.30 GiB0.60 GiB3.71 GiB0.01 GiB21±37%
zeta-2.1I1-IQ2_XXS8.3B2.34 GiB0.53 GiB3.71 GiB0.01 GiB23±30%
Holo-3.1-4BI1-IQ4_NL5.2B2.76 GiB0.13 GiB3.71 GiB0.01 GiB23±30%
AfriqueQwen3.5-4BI1-IQ4_NL5.2B2.76 GiB0.13 GiB3.71 GiB0.01 GiB23±30%
TimeOmni-1-4BI1-IQ4_NL5.2B2.76 GiB0.13 GiB3.71 GiB0.01 GiB23±30%
Hubble-4B-v1IQ4_XS4.5B2.36 GiB0.53 GiB3.71 GiB0.01 GiB23±30%
Aura-4BI1-IQ4_XS4.5B2.36 GiB0.53 GiB3.71 GiB0.01 GiB23±30%
magnum-v2-4bI1-IQ4_XS4.5B2.36 GiB0.53 GiB3.71 GiB0.01 GiB23±30%
Impish_LLAMA_4BIQ4_XS4.5B2.36 GiB0.53 GiB3.71 GiB0.01 GiB23±30%
Darwin-4B-ChimeraQ4_K_L4.0B2.51 GiB0.38 GiB3.70 GiB0.02 GiB23±30%
Llama-3.1-8B-InstructUD-IQ2_XXS8.0B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
Llama-3.1-Nemotron-Nano-8B-v1UD-IQ2_XXS8.0B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
DeepSeek-R1-Distill-Llama-8BUD-IQ2_XXS8.0B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
MiMo-VL-7B-RLUD-IQ2_XXS8.3B2.28 GiB0.60 GiB3.70 GiB0.02 GiB23±30%
Qwen3-Embedding-4BQ4_K4.0B2.29 GiB0.60 GiB3.70 GiB0.02 GiB23±30%
deepseek-coder-5.7bmqa-baseQ3_K_L5.7B2.81 GiB0.07 GiB3.70 GiB0.02 GiB23±30%
dolphin-2.9.3-mistral-7B-32kI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
Mistral-7B-v0.3IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
MiniCPM-V-4Q6_K4.1B2.76 GiB0.13 GiB3.70 GiB0.02 GiB23±30%
Mistral-7B-Instruct-v0.3-ParasiteI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
Mistral-7B-Instruct-v0.3-JbliteratedI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
Mathstral-7B-v0.1IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
gemma-3n-E2B-itQ4_K_M5.4B2.82 GiB0.07 GiB3.70 GiB0.02 GiB23±30%
SciPhi-Self-RAG-Mistral-7B-32kKV unresolvedI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
dolphin-2.2.1-mistral-7bKV unresolvedI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
OpenChat-3.5-7B-Qwen-v2.0KV unresolvedI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
openchat-3.5-0106KV unresolvedI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
Mistral-7B-Instruct-v0.1KV unresolvedI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
Mistral-7B-Instruct-v0.2I1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
ContextualKunoichi_KTO-7BI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
xLAM-7b-rI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
MegaBeam-Mistral-7B-512kIQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
Ninja-v1-RP-WIPKV unresolvedI1-IQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
Kunoichi-DPO-v2-7BKV unresolvedIQ2_M7.2B2.33 GiB0.53 GiB3.70 GiB0.02 GiB23±30%
umt5-xxlQ3_K_M5.7B2.85 GiB0.00 GiB3.70 GiB0.02 GiB23±30%
Phi-4-mini-reasoningQ4_13.8B2.35 GiB0.53 GiB3.69 GiB0.03 GiB23±30%
Phi-4-mini-instructQ4_13.8B2.35 GiB0.53 GiB3.69 GiB0.03 GiB23±30%
whisper-mediumF32764M2.85 GiB0.00 GiB3.69 GiB0.03 GiB23±30%
whisper-medium.enF32764M2.85 GiB0.00 GiB3.69 GiB0.03 GiB23±30%
granite-4.0-h-tinyMoEQ3_K_S6.9B2.89 GiB0.03 GiB3.69 GiB0.03 GiB73±37%
Teuken-7B-instruct-research-v0.4I1-IQ2_XXS7.5B2.72 GiB0.13 GiB3.69 GiB0.03 GiB23±30%
OLMoE-1B-7B-0924-InstructMoEI1-Q2_K6.9B2.39 GiB0.53 GiB3.69 GiB0.03 GiB38±37%
Llama-3.2-3B-Instruct-abliteratedI1-Q5_K_M3.6B2.41 GiB0.46 GiB3.69 GiB0.03 GiB23±30%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a Arc A310 4GB run?
802 of 2118 indexed open-weight models fit a Arc A310 4GB at 8,192 context with q8_0 KV cache, the largest being ToriiGate-0.5 at IQ4_NL. That covers text, vision-language, image, video and speech models.
How much usable memory does a Arc A310 4GB actually have?
Its nameplate is 4 GB, but about 3.72 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Arc A310 4GB fast for local AI?
Its memory bandwidth is 124 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.