NVIDIA · workstation

RTX A2000

RTX A2000 has 6 GB of VRAM at 288 GB/s — about 5.58 GiB usable after driver and compositor overhead. 931 of 2118 indexed models fit at 32K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
288 GB/s
192-bit bus
Tensor FP16
32 TF
dense
TDP
70 W
$449 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 767embedding 25vision language 80audio tts 19audio asr 38video 2

What fits at 32K context

largest quantization that fits, per model · 931 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Aya-Medikal-V2I1-IQ2_XXS8.0B2.41 GiB2.13 GiB5.58 GiB0.00 GiB36±22%
granite-4.0-h-tinyMoEQ5_06.9B4.48 GiB0.13 GiB5.58 GiB0.00 GiB111±37%
granite-4.0-h-tiny-baseMoEQ5_06.9B4.48 GiB0.13 GiB5.58 GiB0.00 GiB111±37%
Marco-Nano-InstructMoEI1-IQ2_XXS8.0B2.75 GiB1.86 GiB5.58 GiB0.00 GiB44±37%
OLMoE-1B-7B-0924-InstructMoEQ2_K_L6.9B2.48 GiB2.13 GiB5.58 GiB0.00 GiB36±37%
nomic-embed-codeQ3_K_L7.1B3.59 GiB0.93 GiB5.57 GiB0.01 GiB36±22%
SmolLM2-1.7B-Instruct-UncensoredQ6_K1.8B1.39 GiB3.19 GiB5.57 GiB0.01 GiB36±22%
Nanbeige4.2-3BUD-Q6_K4.2B3.09 GiB1.46 GiB5.57 GiB0.01 GiB36±22%
Apertus-8B-Instruct-2509UD-IQ2_XXS8.1B2.38 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
GLM-4.1V-9B-ThinkingQ2_K_L10.3B3.87 GiB0.66 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-9BUD-IQ3_XXS9.7B4.00 GiB0.53 GiB5.57 GiB0.01 GiB36±22%
salamandra-7b-instruct-2606I1-IQ2_XXS7.8B2.41 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
qwen-indic-v1I1-IQ2_XXS7.6B2.13 GiB2.39 GiB5.55 GiB0.03 GiB36±22%
orpheus-3b-0.1-pretrainedQ5_13.8B2.68 GiB1.86 GiB5.55 GiB0.03 GiB36±22%
Fara1.5-9BIQ3_XXS9.4B3.98 GiB0.53 GiB5.55 GiB0.03 GiB36±22%
QwenPaw-Flash-9BIQ3_XXS9.4B3.98 GiB0.53 GiB5.55 GiB0.03 GiB36±22%
grug-9bIQ3_XXS9.4B3.98 GiB0.53 GiB5.55 GiB0.03 GiB36±22%
OmniCoder-9BIQ3_XXS9.4B3.98 GiB0.53 GiB5.55 GiB0.03 GiB36±22%
Ornith-1.0-9BIQ3_XXS9.2B3.98 GiB0.53 GiB5.55 GiB0.03 GiB36±22%
Qwen3.5-9B-NeoIQ3_XXS9.7B3.98 GiB0.53 GiB5.55 GiB0.03 GiB36±22%
Qwen3-VL-8B-InstructUD-IQ1_S8.8B2.13 GiB2.39 GiB5.55 GiB0.03 GiB36±22%
Ministral-3-8B-Instruct-2512UD-IQ1_M8.9B2.25 GiB2.26 GiB5.54 GiB0.04 GiB36±22%
Qwen3-VL-8B-ThinkingUD-IQ1_S8.8B2.12 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
Qwen3-8BUD-IQ1_S8.2B2.12 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
GrammarCoder-7B-BaseI1-Q3_K_M7.6B3.56 GiB0.93 GiB5.54 GiB0.04 GiB36±22%
internlm3-8b-instructQ3_K_S8.8B3.72 GiB0.80 GiB5.54 GiB0.04 GiB36±22%
Crow-9B-HERETIC-4.6I1-IQ3_S9.4B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKINGI1-IQ3_S9.4B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSOREDI1-IQ3_S9.4B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Qwen3.5-9B-Claude-4.6-OS-HERETIC-UNCENSORED-INSTRUCTI1-IQ3_S9.4B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSOREDI1-IQ3_S9.4B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
NaNovel-9BI1-IQ3_S9.7B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Qwen3.5-9B-Unredacted-MAXI1-IQ3_S9.4B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Qwen3.5-9B-abliteratedI1-IQ3_S9.4B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Ken3.5-9BI1-IQ3_S9.7B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Qwen3.5-9B-BaseIQ3_S9.7B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Qwen3.5-9B-gemini-3.1-opus-4.6-reasoningI1-IQ3_S9.4B3.97 GiB0.53 GiB5.54 GiB0.04 GiB36±22%
Ministral-3-8B-Reasoning-2512UD-IQ1_M8.9B2.24 GiB2.26 GiB5.54 GiB0.04 GiB36±22%
DeepSeek-R1-0528-Qwen3-8BUD-IQ1_S8.2B2.11 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
Parable-Granite-4.1-8B-Claude-Fable-5I1-IQ1_S8.4B1.85 GiB2.66 GiB5.54 GiB0.04 GiB36±22%
Vero-Qwen35-9B-BaseI1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Vero-Qwen35-9BI1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-ThinkingI1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Morphos-9BI1-Q3_K_S9.0B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwable-9B-Claude-Fable-5-hereticI1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Holo-3.1-9BI1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwable-9B-Claude-Fable-5I1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwen3.5-9B-imabari-v2I1-Q3_K_S9.7B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwen3.5-9B-abliterated-v2-MAXI1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
OmniCoder-9B-Claude-Opus-High-Reasoning-DistillI1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwable-9B-Claude-Fable-5-StraTAI1-Q3_K_S9.0B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwable-9B-Claude-Fable-5-OBLITERATEDI1-Q3_K_S9.0B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwen3.5-9B-RpRMax-v1I1-Q3_K_S9.7B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
AdQWENistrator-9BI1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
cajal-9b-v2-fullI1-Q3_K_S9.0B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwen3.5-9B-ultra-uncensored-hereticQ3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Holo-3.1-9B-CoderI1-Q3_K_S9.0B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
PlutoI1-Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Holo-3.1-9B-abliterated-rdoI1-Q3_K_S9.0B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
Qwen3.5-9B-Uncensored-cyber-v3Q3_K_S9.4B3.97 GiB0.53 GiB5.53 GiB0.05 GiB36±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX A2000 run?
931 of 2118 indexed open-weight models fit a RTX A2000 at 32,768 context with q8_0 KV cache, the largest being Aya-Medikal-V2 at I1-IQ2_XXS. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A2000 actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A2000 fast for local AI?
Its memory bandwidth is 288 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.