RTX A2000
RTX A2000 has 12 GB of VRAM at 288 GB/s — about 11.16 GiB usable after driver and compositor overhead. 972 of 2118 indexed models fit at 64K context with f16 KV.
What fits at 64K context
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| Ling-liteMoE | Q2_K_L | 16.8B | 6.67 GiB | 3.50 GiB | 11.16 GiB | 0.00 GiB | 21±37% |
| rnj-1-instruct | UD-IQ1_M | 8.3B | 2.11 GiB | 8.00 GiB | 11.16 GiB | 0.00 GiB | 16±22% |
| zeta-2.1 | I1-IQ1_M | 8.3B | 2.12 GiB | 8.00 GiB | 11.16 GiB | 0.00 GiB | 16±22% |
| Hubble-4B-v1 | Q3_K_M | 4.5B | 2.14 GiB | 8.00 GiB | 11.15 GiB | 0.01 GiB | 16±22% |
| Aura-4B | I1-Q3_K_M | 4.5B | 2.14 GiB | 8.00 GiB | 11.15 GiB | 0.01 GiB | 16±22% |
| magnum-v2-4b | I1-Q3_K_M | 4.5B | 2.14 GiB | 8.00 GiB | 11.15 GiB | 0.01 GiB | 16±22% |
| Impish_LLAMA_4B | Q3_K_M | 4.5B | 2.14 GiB | 8.00 GiB | 11.15 GiB | 0.01 GiB | 16±22% |
| Llama-3.1-Minitron-4B-Width-Base | Q3_K_M | 4.5B | 2.14 GiB | 8.00 GiB | 11.15 GiB | 0.01 GiB | 16±22% |
| Apertus-8B-Instruct-2509 | UD-IQ1_S | 8.1B | 2.08 GiB | 8.00 GiB | 11.14 GiB | 0.02 GiB | 16±22% |
| Qwen3.6-35B-A3B-REAM-160-ru-agentMoE | IQ3_XXS | 23.6B | 8.89 GiB | 1.25 GiB | 11.14 GiB | 0.02 GiB | 42±37% |
| OLMoE-1B-7B-0924-InstructMoE | I1-IQ2_M | 6.9B | 2.17 GiB | 8.00 GiB | 11.14 GiB | 0.02 GiB | 11±37% |
| Voxtral-Mini-3B-2507 | Q5_K_S | 4.7B | 2.63 GiB | 7.50 GiB | 11.14 GiB | 0.02 GiB | 16±22% |
| Kimi-VL-A3B-Thinking-2506MoE | IQ4_XS | 16.4B | 8.21 GiB | 1.90 GiB | 11.13 GiB | 0.03 GiB | 30±37% |
| Kimi-VL-A3B-InstructMoE | IQ4_XS | 16.4B | 8.21 GiB | 1.90 GiB | 11.13 GiB | 0.03 GiB | 30±37% |
| orpheus-3b-0.1-pretrained | Q6_K_L | 3.8B | 3.12 GiB | 7.00 GiB | 11.13 GiB | 0.03 GiB | 16±22% |
| Trinity-MiniMoE | Q2_K | 26.1B | 9.01 GiB | 1.12 GiB | 11.12 GiB | 0.04 GiB | 42±37% |
| Falcon3-7B-Instruct | Q2_K_L | 7.5B | 3.05 GiB | 7.00 GiB | 11.12 GiB | 0.04 GiB | 16±22% |
| VibeVoice-1.5B | F32 | 2.7B | 10.07 GiB | 0.00 GiB | 11.12 GiB | 0.04 GiB | 16±22% |
| Qwen3-16B-A3BMoE | IQ2_XXS | 16.0B | 4.12 GiB | 6.00 GiB | 11.12 GiB | 0.04 GiB | 14±37% |
| granite-3.1-3b-a800m-instructMoE | F16 | 3.3B | 6.15 GiB | 4.00 GiB | 11.12 GiB | 0.04 GiB | 17±37% |
| granite-34b-code-base-8k | I1-IQ2_S | 33.7B | 10.04 GiB | 0.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| CycleGRPO-4B | I1-IQ1_S | 4.8B | 1.10 GiB | 9.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Hunyuan-7B-Instruct | IQ2_XXS | 7.5B | 2.07 GiB | 8.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| gemma-3-12b-it-vl-Gemini-3-Pro-Preview-Heretic-Uncensored-Thinking | I1-Q3_K_M | 12.2B | 5.60 GiB | 4.47 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-Thinking | I1-Q3_K_M | 12.2B | 5.60 GiB | 4.47 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| gemma-3-12b-it-ultra-uncensored-heretic | Q3_K_M | 12.2B | 5.60 GiB | 4.47 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| gemma-3-12b-it-vl-GLM-4.7-Flash-Heretic-Uncensored-Thinking | I1-Q3_K_M | 12.2B | 5.60 GiB | 4.47 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Floppa-12B-Gemma3-Uncensored | I1-Q3_K_M | 12.2B | 5.60 GiB | 4.47 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| gemma-3-12b-it-heretic | I1-Q3_K_M | 12.2B | 5.60 GiB | 4.47 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| gemma-3-12b-it-abliterated | Q3_K_M | 12.2B | 5.60 GiB | 4.47 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| gemma-3-12b-it-abliterated-v2 | Q3_K_M | 11.8B | 5.60 GiB | 4.47 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| gemma-3-12b-it | Q3_K_M | 12.2B | 5.60 GiB | 4.47 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Marco-Nano-InstructMoE | I1-Q2_K_S | 8.0B | 3.13 GiB | 7.00 GiB | 11.11 GiB | 0.05 GiB | 13±37% |
| Phi-4-mini-instruct-abliterated | Q3_K_L | 3.8B | 2.10 GiB | 8.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Phi-4-mini-reasoning | Q3_K_L | 3.8B | 2.10 GiB | 8.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| Phi-4-mini-instruct | Q3_K_L | 3.8B | 2.10 GiB | 8.00 GiB | 11.11 GiB | 0.05 GiB | 16±22% |
| legitus-instruct-v1 | I1-IQ1_M | 8.1B | 2.04 GiB | 8.00 GiB | 11.10 GiB | 0.06 GiB | 16±22% |
| Aya-Medikal-V2 | I1-IQ1_S | 8.0B | 2.06 GiB | 8.00 GiB | 11.10 GiB | 0.06 GiB | 16±22% |
| ERNIE-4.5-21B-A3B-PT | UD-IQ1_S | 21.9B | 6.58 GiB | 3.50 GiB | 11.10 GiB | 0.06 GiB | 16±22% |
| Carnice-Qwen3.6-MoE-35B-A3BMoE | I1-IQ2_XXS | 36.0B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Qwen35B-Agent-R2-AbliteratedMoE | I1-IQ2_XXS | 34.7B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoE | I1-IQ2_XXS | 36.0B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Darwin-35B-A3B-OpusMoE | I1-IQ2_XXS | 36.0B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Qwen35B-Agent-R2MoE | I1-IQ2_XXS | 34.7B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Carnice-MoE-35B-A3BMoE | I1-IQ2_XXS | 36.0B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| spoomplesmaxx-flash-35B-A3MoE | I1-IQ2_XXS | 35.1B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Huihui-Qwen3.6-35B-A3B-Claude-4.7-Opus-abliteratedMoE | I1-IQ2_XXS | 36.0B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Qwen3.6-35B-A3B-Uncensored-AggressiveMoE | I1-IQ2_XXS | 35.1B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| WorldSim-Opus-3.6-35B-A3BMoE | I1-IQ2_XXS | 35.1B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Qwen3.6-35B-A3B-abliterated-MAXMoE | I1-IQ2_XXS | 35.1B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Huihui-Qwen3.6-35B-A3B-abliteratedMoE | I1-IQ2_XXS | 36.0B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Qwopus3.6-35B-A3B-v1MoE | I1-IQ2_XXS | 36.0B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Qwen3.6-35B-A3B-StyleTuneMoE | I1-IQ2_XXS | 35.1B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Qwen3.6-35B-A3B-abliteratedMoE | I1-IQ2_XXS | 35.1B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| 0GM-1.0-35B-A3B-0427MoE | I1-IQ2_XXS | 36.0B | 8.85 GiB | 1.25 GiB | 11.10 GiB | 0.06 GiB | 45±37% |
| Wan2.2-Distill-Models | Q5_K_M | 14.3B | 10.06 GiB | 0.00 GiB | 11.09 GiB | 0.07 GiB | 16±22% |
| SkyReels-V2-DF-14B-540P | Q5_K_M | 14.3B | 10.06 GiB | 0.00 GiB | 11.09 GiB | 0.07 GiB | 16±22% |
| dolphin-2.9.3-mistral-7B-32k | I1-IQ2_XS | 7.2B | 2.05 GiB | 8.00 GiB | 11.09 GiB | 0.07 GiB | 16±22% |
| Mistral-7B-v0.3 | IQ2_XS | 7.2B | 2.05 GiB | 8.00 GiB | 11.09 GiB | 0.07 GiB | 16±22% |
| Mistral-7B-Instruct-v0.3-Parasite | I1-IQ2_XS | 7.2B | 2.05 GiB | 8.00 GiB | 11.09 GiB | 0.07 GiB | 16±22% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Measured on this card
| Workload◍ | Median | Middle 50% | Runs |
|---|---|---|---|
| Image generation | 4.98 it/s | 3.58–6.36 | 66 |
Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.
Questions people ask
- What AI models can a RTX A2000 run?
- 972 of 2118 indexed open-weight models fit a RTX A2000 at 65,536 context with f16 KV cache, the largest being Ling-lite at Q2_K_L. That covers text, vision-language, image, video and speech models.
- How much usable memory does a RTX A2000 actually have?
- Its nameplate is 12 GB, but about 11.16 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a RTX A2000 fast for local AI?
- Its memory bandwidth is 288 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.