NVIDIA · workstation

RTX A4000

RTX A4000 has 16 GB of VRAM at 448 GB/s — about 14.88 GiB usable after driver and compositor overhead. 1680 of 2118 indexed models fit at 64K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
GDDR6
Bandwidth
448 GB/s
256-bit bus
Tensor FP16
77 TF
dense
TDP
140 W
$1000 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1427vision language 151audio asr 39video 15embedding 26audio tts 21image 1

What fits at 64K context

largest quantization that fits, per model · 1680 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Tess-3-Mistral-Nemo-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Vikhr-Nemo-12B-Instruct-R-21-09-24Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Wayfarer-2-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Wayfarer-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
writing-roleplay-20k-context-nemo-12b-v1.0Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Dans-PersonalityEngine-V1.3.0-12bQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Lumimaid-v0.2-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Lumimaid-Magnum-v4-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Captain-Eris_Violet-V0.420-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Mistral-Nemo-Gutenberg-Doppel-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Mistral-Nemo-Instruct-2407Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
MN-12b-RP-InkQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
magnum-v4-12bQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Mistral-Nemo-12B-ArliAI-RPMax-v1.1Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Rocinante-X-12B-v1Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Dans-SakuraKaze-V1.0.0-12bQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
MN-12B-Mag-Mell-R1Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
pixtral-12bQ5_K_L12.7B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Crimson_Dawn-v0.2Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Magnum-Picaro-0.7-v2-12bQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Nera_Noctis-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Mistral-Nemo-Prism-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Chronos-Gold-12B-1.0Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
MN-12B-Celeste-V1.9Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
magnum-v2.5-12b-ktoQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
magnum-v2-12bQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
NemoMix-Unleashed-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Rocinante-12B-v1.1Q5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
BlackSheep-RP-12BQ5_K_L12.2B8.51 GiB5.31 GiB14.87 GiB0.01 GiB19±22%
Rocinante-XL-16B-v1IQ3_XS16.1B6.65 GiB7.17 GiB14.87 GiB0.01 GiB19±22%
Ling-liteMoEQ5_K_L16.8B12.02 GiB1.86 GiB14.87 GiB0.01 GiB40±37%
granite-20b-code-instruct-8kQ5_K_M20.1B13.79 GiB0.00 GiB14.87 GiB0.01 GiB19±22%
granite-20b-code-base-8kI1-Q5_K_M20.1B13.79 GiB0.00 GiB14.87 GiB0.01 GiB19±22%
granite-34b-code-base-8kI1-IQ3_S33.7B13.79 GiB0.00 GiB14.87 GiB0.01 GiB19±22%
Skywork-R1V3-38BIQ3_M38.4B13.79 GiB0.00 GiB14.86 GiB0.02 GiB19±22%
Devstral-Small-2507Q2_K_L23.6B8.43 GiB5.31 GiB14.86 GiB0.02 GiB19±22%
Magistral-Small-2509Q2_K_L24.0B8.43 GiB5.31 GiB14.86 GiB0.02 GiB19±22%
Magistral-Small-2507Q2_K_L23.6B8.43 GiB5.31 GiB14.86 GiB0.02 GiB19±22%
diffusiongemma-26B-A4B-it-HERETIC-UncensoredMoEQ3_K_M25.8B12.38 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
diffusiongemma-26B-A4B-itMoEQ3_K_M25.8B12.38 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
Olmo-3.1-32B-InstructQ2_K32.2B11.18 GiB2.57 GiB14.85 GiB0.03 GiB19±22%
Olmo-3.1-32B-ThinkQ2_K32.2B11.18 GiB2.57 GiB14.85 GiB0.03 GiB19±22%
Olmo-3-32B-ThinkQ2_K32.2B11.18 GiB2.57 GiB14.85 GiB0.03 GiB19±22%
Apriel-1.6-15b-ThinkerIQ4_XS14.9B7.43 GiB6.38 GiB14.85 GiB0.03 GiB19±22%
Magistral-Small-2509-VisionQ2_K_S24.0B8.42 GiB5.31 GiB14.85 GiB0.03 GiB19±22%
WizardCoder-Python-34B-V1.0I1-IQ1_M33.7B7.38 GiB6.38 GiB14.85 GiB0.03 GiB19±22%
Phind-CodeLlama-34B-Python-v1I1-IQ1_M33.7B7.38 GiB6.38 GiB14.85 GiB0.03 GiB19±22%
Phind-CodeLlama-34B-v2I1-IQ1_M33.7B7.38 GiB6.38 GiB14.85 GiB0.03 GiB19±22%
dolphin-2.6-mistral-7bQ5_K_M7.2B9.56 GiB4.25 GiB14.85 GiB0.03 GiB19±22%
DeepSeek-R1-Distill-Llama-8B-AbliteratedI1-Q4_18.0B9.56 GiB4.25 GiB14.85 GiB0.03 GiB19±22%
Goetia-26B-A4B-v1.3-Absolute-Heretic-ARAMoEI1-Q3_K_M25.8B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
Frank-26B-A4BMoEI1-Q3_K_M26.5B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
G4-MeroMero-26B-A4B-it-uncensored-hereticMoEI1-Q3_K_M25.8B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
EVE-26b-XENO-HATMoEI1-Q3_K_M25.8B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
Gemma-4-26B-A4B-Animus-V14.1-FFT-hereticMoEI1-Q3_K_M25.8B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
gemma-4-26B-A4B-it-qat-q4_0-unquantized-hereticMoEI1-Q3_K_M25.8B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
gemma-4-26B-A4B-it-Claude-Opus-DistillMoEQ3_K_M26.5B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
Huihui-gemma-4-26B-A4B-it-qat-q4_0-unquantized-abliteratedMoEI1-Q3_K_M26.5B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
G4-MeroMero-26B-A4BMoEI1-Q3_K_M25.8B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
gemma-4-26B-A4B-it-Claude-Opus-Distill-v2MoEQ3_K_M26.5B12.37 GiB1.48 GiB14.85 GiB0.03 GiB18±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation12.43 it/s9.8214.54304
Prompt processing2452.65 tok/s2018.102695.4112
Text generation81.90 tok/s78.4483.7310
Benchmarked· n=304

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a RTX A4000 run?
1680 of 2118 indexed open-weight models fit a RTX A4000 at 65,536 context with q8_0 KV cache, the largest being Tess-3-Mistral-Nemo-12B at Q5_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A4000 actually have?
Its nameplate is 16 GB, but about 14.88 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A4000 fast for local AI?
Its memory bandwidth is 448 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.