NVIDIA · workstation

RTX A400

RTX A400 has 4 GB of VRAM at 96 GB/s — about 3.72 GiB usable after driver and compositor overhead. 578 of 2118 indexed models fit at 16K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
4 GB
GDDR6
Bandwidth
96 GB/s
64-bit bus
Tensor FP16
11 TF
dense
TDP
50 W
$135 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 454audio tts 19audio asr 36vision language 47video 2embedding 20

What fits at 16K context

largest quantization that fits, per model · 578 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
LFM2.5-Audio-1.5B-JPF161.5B2.67 GiB0.00 GiB3.72 GiB0.00 GiB20±22%
FrickFritz-4BI1-Q4_04.7B2.44 GiB0.27 GiB3.72 GiB0.00 GiB20±22%
qwen3.5-4b-agentic-coder-v4I1-Q4_04.7B2.44 GiB0.27 GiB3.72 GiB0.00 GiB20±22%
Myth-4BI1-Q4_04.3B2.44 GiB0.27 GiB3.72 GiB0.00 GiB20±22%
Qwen3.5-4B-UncensoredI1-Q4_04.7B2.44 GiB0.27 GiB3.72 GiB0.00 GiB20±22%
JOSIE-2-4B-PreviewI1-Q4_04.7B2.44 GiB0.27 GiB3.72 GiB0.00 GiB20±22%
Surogate-3.5-4BI1-Q4_05.3B2.44 GiB0.27 GiB3.72 GiB0.00 GiB20±22%
Gemma-3-4b-it-Uncensored-DBL-XI1-IQ4_NL4.7B2.40 GiB0.30 GiB3.71 GiB0.01 GiB20±22%
umt5-xxlQ3_K_S5.7B2.66 GiB0.00 GiB3.71 GiB0.01 GiB21±22%
Amaretto-3BIQ4_XS4.3B1.84 GiB0.86 GiB3.71 GiB0.01 GiB20±22%
orpheus-3b-0.1-ftIQ4_XS3.8B1.77 GiB0.93 GiB3.71 GiB0.01 GiB20±22%
Qwen2.5-3BQ5_13.1B2.40 GiB0.30 GiB3.71 GiB0.01 GiB20±22%
Newton-bot-3-VLM-mini-4BQ4_04.7B2.43 GiB0.27 GiB3.71 GiB0.01 GiB20±22%
Darwin-4B-ChimeraI1-Q4_K_S4.0B2.22 GiB0.48 GiB3.71 GiB0.01 GiB20±22%
Qwen3.5-4B-NSFW-ARA-Heretic-LiteroticaI1-IQ4_NL4.2B2.43 GiB0.27 GiB3.71 GiB0.01 GiB20±22%
Qwen3.5-4B-RpRMax-v1I1-IQ4_NL4.7B2.43 GiB0.27 GiB3.71 GiB0.01 GiB20±22%
Holo-3.1-4B-uncensored-hereticI1-IQ4_NL4.5B2.43 GiB0.27 GiB3.71 GiB0.01 GiB20±22%
GRaPE-2-MiniI1-IQ4_NL4.7B2.43 GiB0.27 GiB3.71 GiB0.01 GiB20±22%
Qwen3.5-DPO-4B-2I1-IQ4_NL4.2B2.43 GiB0.27 GiB3.71 GiB0.01 GiB20±22%
Huihui-Qwen3.5-4B-Claude-4.6-Opus-abliteratedI1-IQ4_NL4.7B2.43 GiB0.27 GiB3.71 GiB0.01 GiB20±22%
Qwopus3.5-4B-v3-hereticI1-IQ4_NL4.5B2.43 GiB0.27 GiB3.71 GiB0.01 GiB20±22%
Aureth-4B-Qwen3.5I1-IQ4_NL4.5B2.43 GiB0.27 GiB3.71 GiB0.01 GiB20±22%
granite-3.1-3b-a800m-instructMoEQ5_K_L3.3B2.21 GiB0.53 GiB3.71 GiB0.01 GiB29±37%
Voxtral-Mini-3B-2507IQ3_XS4.7B1.70 GiB1.00 GiB3.71 GiB0.01 GiB20±22%
Qwen2.5-Omni-7BUD-IQ2_M10.7B2.66 GiB0.00 GiB3.70 GiB0.02 GiB21±22%
SmolLM2-1.7B-InstructQ5_K_S1.7B1.11 GiB1.59 GiB3.70 GiB0.02 GiB20±22%
Felldude-Uncensored-Ministral3-3B-bf16I1-IQ4_XS3.8B1.82 GiB0.86 GiB3.70 GiB0.02 GiB20±22%
Ministral-3-3B-Instruct-2512-BF16IQ4_XS4.3B1.82 GiB0.86 GiB3.70 GiB0.02 GiB20±22%
Ministral-3-3B-Instruct-2512IQ4_XS3.8B1.82 GiB0.86 GiB3.70 GiB0.02 GiB20±22%
Ministral-3-3B-Reasoning-2512IQ4_XS4.3B1.82 GiB0.86 GiB3.70 GiB0.02 GiB20±22%
Vikhr-Gemma-2B-instructQ6_K_L2.6B2.14 GiB0.55 GiB3.70 GiB0.02 GiB20±22%
gemma-2-2b-it-abliteratedQ6_K_L2.6B2.14 GiB0.55 GiB3.70 GiB0.02 GiB20±22%
gemma-2-2b-itQ6_K_L2.6B2.14 GiB0.55 GiB3.70 GiB0.02 GiB20±22%
Gemmasutra-Mini-2B-v1Q6_K_L2.6B2.14 GiB0.55 GiB3.70 GiB0.02 GiB20±22%
Parable-Granite-4.1-3B-Claude-Fable-5I1-Q4_13.4B2.03 GiB0.66 GiB3.70 GiB0.02 GiB20±22%
granite-4.0-microQ4_13.4B2.03 GiB0.66 GiB3.70 GiB0.02 GiB20±22%
granite-4.1-3bQ4_13.4B2.03 GiB0.66 GiB3.70 GiB0.02 GiB20±22%
granite-4.0-micro-baseQ4_13.4B2.03 GiB0.66 GiB3.70 GiB0.02 GiB20±22%
EVA-Yi-1.5-9B-32K-V1I1-IQ1_S8.8B1.88 GiB0.80 GiB3.70 GiB0.02 GiB20±22%
Yi-Coder-9B-ChatIQ1_S8.8B1.88 GiB0.80 GiB3.70 GiB0.02 GiB20±22%
Qwopus3.5-4B-v3IQ4_XS4.7B2.42 GiB0.27 GiB3.69 GiB0.03 GiB20±22%
Luna-7B-A4BMoEI1-IQ1_S6.7B1.49 GiB1.20 GiB3.69 GiB0.03 GiB16±37%
Aura-4BI1-Q2_K_S4.5B1.61 GiB1.06 GiB3.69 GiB0.03 GiB20±22%
moondream2F161.9B2.64 GiB0.00 GiB3.69 GiB0.03 GiB21±22%
Nanbeige4.2-3BIQ3_M4.2B1.94 GiB0.73 GiB3.69 GiB0.03 GiB21±22%
Llama-3.2-3B-Instruct-uncensoredQ2_K_L3.6B1.75 GiB0.93 GiB3.69 GiB0.03 GiB20±22%
LFM2-2.6BQ8_02.6B2.55 GiB0.13 GiB3.69 GiB0.03 GiB20±22%
LFM2-2.6B-TranscriptQ8_02.6B2.55 GiB0.13 GiB3.69 GiB0.03 GiB20±22%
LFM2-VL-3BQ8_03.0B2.55 GiB0.13 GiB3.69 GiB0.03 GiB20±22%
CycleGRPO-4BI1-IQ2_S4.8B1.48 GiB1.20 GiB3.68 GiB0.04 GiB20±22%
Yi-Coder-1.5B-ChatQ5_K_L1.5B1.10 GiB1.59 GiB3.68 GiB0.04 GiB20±22%
Yi-Coder-1.5BQ5_K_L1.5B1.10 GiB1.59 GiB3.68 GiB0.04 GiB20±22%
magnum-v2-4bI1-IQ2_M4.5B1.60 GiB1.06 GiB3.68 GiB0.04 GiB21±22%
Impish_LLAMA_4BIQ2_M4.5B1.60 GiB1.06 GiB3.68 GiB0.04 GiB21±22%
VoxCPM2Q8_02.3B2.63 GiB0.00 GiB3.68 GiB0.04 GiB21±22%
Qwen3.5-4BQ3_K_M4.7B2.40 GiB0.27 GiB3.68 GiB0.04 GiB21±22%
Dolphin3.0-Qwen2.5-3bQ6_K3.1B2.36 GiB0.30 GiB3.67 GiB0.05 GiB21±22%
Qwen2.5-Coder-3B-Instruct-abliteratedI1-Q6_K3.1B2.36 GiB0.30 GiB3.67 GiB0.05 GiB21±22%
GRM-Kerlin-3b-AbliteratedI1-Q6_K3.1B2.36 GiB0.30 GiB3.67 GiB0.05 GiB21±22%
Mythos-nanoI1-Q6_K3.1B2.36 GiB0.30 GiB3.67 GiB0.05 GiB21±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a RTX A400 run?
578 of 2118 indexed open-weight models fit a RTX A400 at 16,384 context with q8_0 KV cache, the largest being LFM2.5-Audio-1.5B-JP at F16. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A400 actually have?
Its nameplate is 4 GB, but about 3.72 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A400 fast for local AI?
Its memory bandwidth is 96 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.