NVIDIA · workstation
RTX A2000
RTX A2000 has 6 GB of VRAM at 288 GB/s — about 5.58 GiB usable after driver and compositor overhead. 931 of 2118 indexed models fit at 32K context with q8_0 KV.
Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
288 GB/s
192-bit bus
Tensor FP16
32 TF
dense
TDP
70 W
$449 MSRP
text 767embedding 25vision language 80audio tts 19audio asr 38video 2
What fits at 32K context
largest quantization that fits, per model · 931 of 2118 indexed
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| Aya-Medikal-V2 | I1-IQ2_XXS | 8.0B | 2.41 GiB | 2.13 GiB | 5.58 GiB | 0.00 GiB | 36±22% |
| granite-4.0-h-tinyMoE | Q5_0 | 6.9B | 4.48 GiB | 0.13 GiB | 5.58 GiB | 0.00 GiB | 111±37% |
| granite-4.0-h-tiny-baseMoE | Q5_0 | 6.9B | 4.48 GiB | 0.13 GiB | 5.58 GiB | 0.00 GiB | 111±37% |
| Marco-Nano-InstructMoE | I1-IQ2_XXS | 8.0B | 2.75 GiB | 1.86 GiB | 5.58 GiB | 0.00 GiB | 44±37% |
| OLMoE-1B-7B-0924-InstructMoE | Q2_K_L | 6.9B | 2.48 GiB | 2.13 GiB | 5.58 GiB | 0.00 GiB | 36±37% |
| nomic-embed-code | Q3_K_L | 7.1B | 3.59 GiB | 0.93 GiB | 5.57 GiB | 0.01 GiB | 36±22% |
| SmolLM2-1.7B-Instruct-Uncensored | Q6_K | 1.8B | 1.39 GiB | 3.19 GiB | 5.57 GiB | 0.01 GiB | 36±22% |
| Nanbeige4.2-3B | UD-Q6_K | 4.2B | 3.09 GiB | 1.46 GiB | 5.57 GiB | 0.01 GiB | 36±22% |
| Apertus-8B-Instruct-2509 | UD-IQ2_XXS | 8.1B | 2.38 GiB | 2.13 GiB | 5.57 GiB | 0.01 GiB | 36±22% |
| GLM-4.1V-9B-Thinking | Q2_K_L | 10.3B | 3.87 GiB | 0.66 GiB | 5.57 GiB | 0.01 GiB | 36±22% |
| Qwen3.5-9B | UD-IQ3_XXS | 9.7B | 4.00 GiB | 0.53 GiB | 5.57 GiB | 0.01 GiB | 36±22% |
| salamandra-7b-instruct-2606 | I1-IQ2_XXS | 7.8B | 2.41 GiB | 2.13 GiB | 5.57 GiB | 0.01 GiB | 36±22% |
| qwen-indic-v1 | I1-IQ2_XXS | 7.6B | 2.13 GiB | 2.39 GiB | 5.55 GiB | 0.03 GiB | 36±22% |
| orpheus-3b-0.1-pretrained | Q5_1 | 3.8B | 2.68 GiB | 1.86 GiB | 5.55 GiB | 0.03 GiB | 36±22% |
| Fara1.5-9B | IQ3_XXS | 9.4B | 3.98 GiB | 0.53 GiB | 5.55 GiB | 0.03 GiB | 36±22% |
| QwenPaw-Flash-9B | IQ3_XXS | 9.4B | 3.98 GiB | 0.53 GiB | 5.55 GiB | 0.03 GiB | 36±22% |
| grug-9b | IQ3_XXS | 9.4B | 3.98 GiB | 0.53 GiB | 5.55 GiB | 0.03 GiB | 36±22% |
| OmniCoder-9B | IQ3_XXS | 9.4B | 3.98 GiB | 0.53 GiB | 5.55 GiB | 0.03 GiB | 36±22% |
| Ornith-1.0-9B | IQ3_XXS | 9.2B | 3.98 GiB | 0.53 GiB | 5.55 GiB | 0.03 GiB | 36±22% |
| Qwen3.5-9B-Neo | IQ3_XXS | 9.7B | 3.98 GiB | 0.53 GiB | 5.55 GiB | 0.03 GiB | 36±22% |
| Qwen3-VL-8B-Instruct | UD-IQ1_S | 8.8B | 2.13 GiB | 2.39 GiB | 5.55 GiB | 0.03 GiB | 36±22% |
| Ministral-3-8B-Instruct-2512 | UD-IQ1_M | 8.9B | 2.25 GiB | 2.26 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3-VL-8B-Thinking | UD-IQ1_S | 8.8B | 2.12 GiB | 2.39 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3-8B | UD-IQ1_S | 8.2B | 2.12 GiB | 2.39 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| GrammarCoder-7B-Base | I1-Q3_K_M | 7.6B | 3.56 GiB | 0.93 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| internlm3-8b-instruct | Q3_K_S | 8.8B | 3.72 GiB | 0.80 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Crow-9B-HERETIC-4.6 | I1-IQ3_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING | I1-IQ3_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSORED | I1-IQ3_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3.5-9B-Claude-4.6-OS-HERETIC-UNCENSORED-INSTRUCT | I1-IQ3_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSORED | I1-IQ3_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| NaNovel-9B | I1-IQ3_S | 9.7B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3.5-9B-Unredacted-MAX | I1-IQ3_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3.5-9B-abliterated | I1-IQ3_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Ken3.5-9B | I1-IQ3_S | 9.7B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3.5-9B-Base | IQ3_S | 9.7B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Qwen3.5-9B-gemini-3.1-opus-4.6-reasoning | I1-IQ3_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Ministral-3-8B-Reasoning-2512 | UD-IQ1_M | 8.9B | 2.24 GiB | 2.26 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| DeepSeek-R1-0528-Qwen3-8B | UD-IQ1_S | 8.2B | 2.11 GiB | 2.39 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Parable-Granite-4.1-8B-Claude-Fable-5 | I1-IQ1_S | 8.4B | 1.85 GiB | 2.66 GiB | 5.54 GiB | 0.04 GiB | 36±22% |
| Vero-Qwen35-9B-Base | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Vero-Qwen35-9B | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-Thinking | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Morphos-9B | I1-Q3_K_S | 9.0B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwable-9B-Claude-Fable-5-heretic | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Holo-3.1-9B | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwable-9B-Claude-Fable-5 | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwen3.5-9B-imabari-v2 | I1-Q3_K_S | 9.7B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwen3.5-9B-abliterated-v2-MAX | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| OmniCoder-9B-Claude-Opus-High-Reasoning-Distill | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwable-9B-Claude-Fable-5-StraTA | I1-Q3_K_S | 9.0B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwable-9B-Claude-Fable-5-OBLITERATED | I1-Q3_K_S | 9.0B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwen3.5-9B-RpRMax-v1 | I1-Q3_K_S | 9.7B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| AdQWENistrator-9B | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| cajal-9b-v2-full | I1-Q3_K_S | 9.0B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwen3.5-9B-ultra-uncensored-heretic | Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Holo-3.1-9B-Coder | I1-Q3_K_S | 9.0B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Pluto | I1-Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Holo-3.1-9B-abliterated-rdo | I1-Q3_K_S | 9.0B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
| Qwen3.5-9B-Uncensored-cyber-v3 | Q3_K_S | 9.4B | 3.97 GiB | 0.53 GiB | 5.53 GiB | 0.05 GiB | 36±22% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Questions people ask
- What AI models can a RTX A2000 run?
- 931 of 2118 indexed open-weight models fit a RTX A2000 at 32,768 context with q8_0 KV cache, the largest being Aya-Medikal-V2 at I1-IQ2_XXS. That covers text, vision-language, image, video and speech models.
- How much usable memory does a RTX A2000 actually have?
- Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a RTX A2000 fast for local AI?
- Its memory bandwidth is 288 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.