NVIDIA · workstation

RTX A2000

RTX A2000 has 6 GB of VRAM at 288 GB/s — about 5.58 GiB usable after driver and compositor overhead. 407 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
288 GB/s
192-bit bus
Tensor FP16
32 TF
dense
TDP
70 W
$449 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 304vision language 41audio asr 29embedding 14audio tts 17video 2

What fits at 128K context

largest quantization that fits, per model · 407 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
gemma-4-E4B-itQ3_K_S8.0B3.60 GiB0.97 GiB5.58 GiB0.00 GiB36±22%
gemma-4-E4B-itQ3_K_S8.0B3.60 GiB0.97 GiB5.58 GiB0.00 GiB36±22%
Darwin-4B-ChimeraI1-Q5_K_M4.0B2.69 GiB1.87 GiB5.58 GiB0.00 GiB36±22%
FrickFritz-4BI1-Q4_04.7B2.44 GiB2.13 GiB5.58 GiB0.00 GiB36±22%
qwen3.5-4b-agentic-coder-v4I1-Q4_04.7B2.44 GiB2.13 GiB5.58 GiB0.00 GiB36±22%
Myth-4BI1-Q4_04.3B2.44 GiB2.13 GiB5.58 GiB0.00 GiB36±22%
Qwen3.5-4B-UncensoredI1-Q4_04.7B2.44 GiB2.13 GiB5.58 GiB0.00 GiB36±22%
JOSIE-2-4B-PreviewI1-Q4_04.7B2.44 GiB2.13 GiB5.58 GiB0.00 GiB36±22%
Surogate-3.5-4BI1-Q4_05.3B2.44 GiB2.13 GiB5.58 GiB0.00 GiB36±22%
EXAONE-4.0-1.2B-abliteratedI1-IQ3_XXS1.5B0.61 GiB3.98 GiB5.57 GiB0.01 GiB35±22%
Newton-bot-3-VLM-mini-4BQ4_04.7B2.43 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
GLM-OCRI1-Q2_K1.3B0.34 GiB4.25 GiB5.57 GiB0.01 GiB35±22%
Qwen3.5-4B-NSFW-ARA-Heretic-LiteroticaI1-IQ4_NL4.2B2.43 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-4B-RpRMax-v1I1-IQ4_NL4.7B2.43 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
Holo-3.1-4B-uncensored-hereticI1-IQ4_NL4.5B2.43 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
GRaPE-2-MiniI1-IQ4_NL4.7B2.43 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
Qwen3.5-DPO-4B-2I1-IQ4_NL4.2B2.43 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
Huihui-Qwen3.5-4B-Claude-4.6-Opus-abliteratedI1-IQ4_NL4.7B2.43 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
Qwopus3.5-4B-v3-hereticI1-IQ4_NL4.5B2.43 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
Aureth-4B-Qwen3.5I1-IQ4_NL4.5B2.43 GiB2.13 GiB5.57 GiB0.01 GiB36±22%
medgemma-4b-itQ6_K_L4.3B3.12 GiB1.42 GiB5.56 GiB0.02 GiB36±22%
amoral-gemma3-4B-v1Q6_K_L4.3B3.12 GiB1.42 GiB5.56 GiB0.02 GiB36±22%
gemma-3-4b-it-abliteratedQ6_K_L4.3B3.12 GiB1.42 GiB5.56 GiB0.02 GiB36±22%
gemma-3-4b-itQ6_K_L4.3B3.12 GiB1.42 GiB5.56 GiB0.02 GiB36±22%
gemma-4-E4B-it-hereticQ3_K_S8.0B3.58 GiB0.97 GiB5.56 GiB0.02 GiB36±22%
Qwopus3.5-4B-v3IQ4_XS4.7B2.42 GiB2.13 GiB5.55 GiB0.03 GiB36±22%
Vikhr-Gemma-2B-instructIQ2_S2.6B0.96 GiB3.57 GiB5.55 GiB0.03 GiB36±22%
Gemmasutra-Mini-2B-v1I1-IQ2_S2.6B0.96 GiB3.57 GiB5.55 GiB0.03 GiB36±22%
Dolphin3.0-Qwen2.5-3bQ5_K_L3.1B2.14 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
Qwen2.5-Coder-3B-Instruct-abliteratedQ5_K_L3.1B2.14 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
Qwen2.5-Coder-3B-InstructQ5_K_L3.1B2.14 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
Qwen2.5-3B-InstructQ5_K_L3.1B2.14 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
Qwen2.5-Coder-3BQ5_K_L3.1B2.14 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
raspberry-3BQ5_K_L3.1B2.14 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
Qwen2.5-3BQ5_K_L3.1B2.14 GiB2.39 GiB5.54 GiB0.04 GiB36±22%
MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_baseI1-IQ2_XXS8.1B2.65 GiB1.86 GiB5.54 GiB0.04 GiB36±22%
Qwen3.5-4BQ3_K_M4.7B2.40 GiB2.13 GiB5.54 GiB0.04 GiB36±22%
Gemma-3-4b-it-Uncensored-DBL-XI1-Q5_K_S4.7B2.82 GiB1.69 GiB5.53 GiB0.05 GiB36±22%
Qwen3.5-4B-BaseQ4_K_S4.7B2.39 GiB2.13 GiB5.52 GiB0.06 GiB36±22%
MiniCPM-V-4Q5_K_M4.1B2.39 GiB2.13 GiB5.52 GiB0.06 GiB36±22%
gemma-2-2b-itIQ2_XS2.6B0.93 GiB3.57 GiB5.52 GiB0.06 GiB36±22%
Qwen3.5-4B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKINGI1-Q4_K_S4.5B2.38 GiB2.13 GiB5.52 GiB0.06 GiB36±22%
Qwen3.5-4B-SOMPOA-heresy-v2I1-Q4_K_S4.5B2.38 GiB2.13 GiB5.52 GiB0.06 GiB36±22%
Qwen3.5-4B-SOMPOA-heresyI1-Q4_K_S4.5B2.38 GiB2.13 GiB5.52 GiB0.06 GiB36±22%
Qwen3.5-4B-Safety-ThinkingI1-Q4_K_S4.2B2.38 GiB2.13 GiB5.52 GiB0.06 GiB36±22%
Huihui-Qwen3.5-4B-abliteratedI1-Q4_K_S4.5B2.38 GiB2.13 GiB5.52 GiB0.06 GiB36±22%
Darkidol-Ballad-4BI1-Q4_K_S4.5B2.38 GiB2.13 GiB5.52 GiB0.06 GiB36±22%
Qwen3.5-4BQ4_K_S4.7B2.38 GiB2.13 GiB5.52 GiB0.06 GiB36±22%
InternVL3_5-8BQ4_K_S8.5B4.47 GiB0.00 GiB5.52 GiB0.06 GiB36±22%
LFM2-8B-A1BMoEQ3_K_M8.3B3.72 GiB0.80 GiB5.51 GiB0.07 GiB63±37%
gemma-4-E4B-it-abliteratedI1-IQ2_M8.0B3.53 GiB0.97 GiB5.51 GiB0.07 GiB36±22%
gemma-4-E4B-uncensoredI1-IQ2_M7.9B3.53 GiB0.97 GiB5.51 GiB0.07 GiB36±22%
gemma-4-E4B-it-qat-q4_0-unquantized-hereticI1-IQ2_M7.9B3.53 GiB0.97 GiB5.51 GiB0.07 GiB36±22%
gemma-4-E4B-it-qat-heretic_decensoredI1-IQ2_M7.9B3.53 GiB0.97 GiB5.51 GiB0.07 GiB36±22%
gemma-4-E4B-it-QAT-SOMPOA-heresyI1-IQ2_M7.9B3.53 GiB0.97 GiB5.51 GiB0.07 GiB36±22%
gemma4-e4b-mahou-nsfwI1-IQ2_M7.9B3.53 GiB0.97 GiB5.51 GiB0.07 GiB36±22%
gemma-4-E4B-it-mentalchat16kI1-IQ2_M7.9B3.53 GiB0.97 GiB5.51 GiB0.07 GiB36±22%
gemma4-E4B-it-abliteratedI1-IQ2_M7.9B3.53 GiB0.97 GiB5.51 GiB0.07 GiB36±22%
gemma-4-E4B-it-OBLITERATEDI1-IQ2_M8.0B3.53 GiB0.97 GiB5.51 GiB0.07 GiB36±22%
Fara1.5-4BIQ4_XS4.5B2.37 GiB2.13 GiB5.51 GiB0.07 GiB36±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX A2000 run?
407 of 2118 indexed open-weight models fit a RTX A2000 at 131,072 context with q8_0 KV cache, the largest being gemma-4-E4B-it at Q3_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A2000 actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A2000 fast for local AI?
Its memory bandwidth is 288 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.