NVIDIA · workstation

RTX A1000

RTX A1000 has 8 GB of VRAM at 192 GB/s — about 7.44 GiB usable after driver and compositor overhead. 1410 of 2118 indexed models fit at 8K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
GDDR6
Bandwidth
192 GB/s
128-bit bus
Tensor FP16
27 TF
dense
TDP
50 W
$365 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1214vision language 102embedding 26audio tts 21video 7image 2audio asr 38

What fits at 8K context

largest quantization that fits, per model · 1410 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Teuken-7B-instruct-research-v0.4Q6_K_L7.5B6.33 GiB0.07 GiB7.44 GiB0.00 GiB17±22%
Qwen3-Coder-REAP-25B-A3BMoEIQ2_XXS24.9B6.24 GiB0.21 GiB7.44 GiB0.00 GiB59±37%
Nexa-AI-4x4B-InstructMoEI1-IQ4_XS12.1B6.11 GiB0.32 GiB7.44 GiB0.00 GiB18±37%
SuperGemma-4-12b-abliteratedI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-hereticI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-hereticI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-uncensored-hereticI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-qat-q4_0-unquantized-uncensored-hereticQ3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Grug-12BI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Aura-Medium-v1-BF16I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-Esper4I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-GuardpointI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Gemma-4-12B-it-AEON-Abliterated-K4-BF16I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-Tachibana-AgentI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12b-marvin-gutenberg-rp-v2I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12b-crownelius-writerI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliteratedI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12b-asterion-agenticI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Huihui-gemma-4-12B-agentic-fable5-abliteratedI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
g4-12b-it-trismegistusI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma4-12b-it-asimovI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
FabGemmaI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Huihui-gemma-4-12B-it-qat-q4_0-unquantized-abliteratedI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-abliterated-uncensoredI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Gemma-4-12b-it-AbliteratedI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-Queen-it-qat-q4_0-unquantizedI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-heretic_decensoredI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Iris-12B-gemma-4-it-qatI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-coder-fable5-composer2.5-v1I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
G4-Starry-Ocean-12BI1-Q3_K_L11.9B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-QAT-SOMPOA-heresyI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-uncensored-opus4.7-cotI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Gemma4-12B-IT-AbliteratedI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12b-it-uncensoredI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Huihui-gemma-4-12B-it-abliteratedI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-hereticI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Tema_Q-X5-12B-ThinkingI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-coder-fable5-composer2.5-v1-bf16I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
swarm-sovereign-12bI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Gemma-4-12B-OBLITERATEDI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Gemma4-12B-UncensoredI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Serenity-12BI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Dark-PaneI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Reelva-12BI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
G4-Starry-Ocean-12B-hereticI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Iris-12B-v1.3.2I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Semancer-12BI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
Iris-12B-v1.2I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12b-heretic-abliteratedI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12B-it-null-space-abliteratedQ3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12b-marvin-gutenbergI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12b-marvin-v2I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
gemma-4-12b-marvin-v1I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
STARK-WEB-12B-v1.7I1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
STARK-WEB-12BI1-Q3_K_L12.0B6.12 GiB0.27 GiB7.43 GiB0.01 GiB17±22%
UncensoredLM-DeepSeek-R1-Distill-Qwen-14BQ3_K_S14.2B5.98 GiB0.40 GiB7.43 GiB0.01 GiB17±22%
Ling-mini-2.0MoEIQ3_XS16.3B6.35 GiB0.09 GiB7.43 GiB0.01 GiB91±37%
glm-4-9b-chat-1mIQ4_XS9.5B4.98 GiB1.41 GiB7.43 GiB0.01 GiB17±22%
LFM2-8B-A1BMoEQ6_K_L8.3B6.41 GiB0.03 GiB7.43 GiB0.01 GiB53±37%
LocateAnything-3BBF163.8B6.34 GiB0.08 GiB7.43 GiB0.01 GiB17±22%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.75 it/s3.594.057
Benchmarked· n=7

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a RTX A1000 run?
1410 of 2118 indexed open-weight models fit a RTX A1000 at 8,192 context with q4_0 KV cache, the largest being Teuken-7B-instruct-research-v0.4 at Q6_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX A1000 actually have?
Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX A1000 fast for local AI?
Its memory bandwidth is 192 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.