NVIDIA · workstation

RTX 4000 SFF Ada Generation

RTX 4000 SFF Ada Generation has 20 GB of VRAM at 280 GB/s — about 18.60 GiB usable after driver and compositor overhead. 1793 of 2118 indexed models fit at 128K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
20 GB
GDDR6
Bandwidth
280 GB/s
160-bit bus
Tensor FP16
77 TF
dense
TDP
70 W
$1250 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1520vision language 170video 16audio asr 39image 1embedding 26audio tts 21

What fits at 128K context

largest quantization that fits, per model · 1793 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Huihui-Qwen3-Coder-Next-abliteratedMoEI1-IQ1_M79.7B16.76 GiB0.84 GiB18.59 GiB0.01 GiB42±37%
gemma-2-27b-itIQ3_XS27.2B10.76 GiB6.70 GiB18.59 GiB0.01 GiB9±22%
magnum-v4-27bIQ3_XS27.2B10.76 GiB6.70 GiB18.59 GiB0.01 GiB9±22%
Luna-7B-A4BMoEF166.7B12.51 GiB5.06 GiB18.58 GiB0.02 GiB8±37%
Carnice-Qwen3.6-MoE-35B-A3BMoEI1-Q3_K_L36.0B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Qwen35B-Agent-R2-AbliteratedMoEI1-Q3_K_L34.7B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Qwen3.6-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoEI1-Q3_K_L36.0B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Darwin-35B-A3B-OpusMoEI1-Q3_K_L36.0B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Qwen35B-Agent-R2MoEI1-Q3_K_L34.7B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Carnice-MoE-35B-A3BMoEI1-Q3_K_L36.0B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
spoomplesmaxx-flash-35B-A3MoEI1-Q3_K_L35.1B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Huihui-Qwen3.6-35B-A3B-Claude-4.7-Opus-abliteratedMoEI1-Q3_K_L36.0B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Qwen3.6-35B-A3B-Uncensored-AggressiveMoEI1-Q3_K_L35.1B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Holo-3.1-35B-A3BMoEQ3_K_L35.1B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
WorldSim-Opus-3.6-35B-A3BMoEI1-Q3_K_L35.1B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Qwen3.6-35B-A3B-abliterated-MAXMoEI1-Q3_K_L35.1B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Huihui-Qwen3.6-35B-A3B-abliteratedMoEI1-Q3_K_L36.0B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Qwopus3.6-35B-A3B-v1MoEI1-Q3_K_L36.0B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
grug-35b-v2MoEQ3_K_L35.1B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Qwen3.6-35B-A3B-StyleTuneMoEI1-Q3_K_L35.1B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Qwen3.6-35B-A3B-abliteratedMoEI1-Q3_K_L35.1B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliteratedMoEQ3_K_L36.0B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
0GM-1.0-35B-A3B-0427MoEI1-Q3_K_L36.0B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
HopCoder-Mini-35B-A3B-VL36MoEQ3_K_L35.1B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Qwen-AgentWorld-35B-A3BMoEUD-IQ4_NL34.7B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
Ornith-1.0-35BMoEUD-IQ4_NL34.7B16.87 GiB0.70 GiB18.58 GiB0.02 GiB41±37%
gemma-4-E2B-it-Uncensored-MAXF325.1B17.33 GiB0.25 GiB18.57 GiB0.03 GiB9±22%
Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16UD-IQ3_S33.0B17.53 GiB0.00 GiB18.57 GiB0.03 GiB9±22%
Gemma-4-31B-Isometry-RPI1-Q2_K32.7B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
Gemma-4-Dark-Gemistry-31BI1-Q2_K32.7B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
Prosopon-31BI1-Q2_K32.7B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
Gemma-4-Novelist-Eclipse-31BI1-Q2_K32.7B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
Giftige-Blume-31B-v1-StyleSwapI1-Q2_K32.7B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
G4-MeroMero-31B-StyleSwapI1-Q2_K32.7B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
Gemma-4-31B-StyleTune-heretic-araI1-Q2_K32.7B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
Pantheon-Reasoning-31B-1.1I1-Q2_K32.7B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
Gemma-4-31B-StyleTuneI1-Q2_K32.7B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
Barcenas-StyleTune-31B-FableI1-Q2_K32.1B11.53 GiB5.95 GiB18.56 GiB0.04 GiB9±22%
medgemma-27b-itIQ4_NL28.8B14.50 GiB2.98 GiB18.56 GiB0.04 GiB9±22%
gemma-3-27b-it-abliteratedIQ4_NL27.4B14.50 GiB2.98 GiB18.56 GiB0.04 GiB9±22%
gemma-3-27b-itIQ4_NL27.4B14.50 GiB2.98 GiB18.56 GiB0.04 GiB9±22%
medgemma-27b-text-itIQ4_NL27.0B14.50 GiB2.98 GiB18.56 GiB0.04 GiB9±22%
WizardCoder-Python-34B-V1.0I1-IQ2_M33.7B10.72 GiB6.75 GiB18.56 GiB0.04 GiB9±22%
Phind-CodeLlama-34B-Python-v1I1-IQ2_M33.7B10.72 GiB6.75 GiB18.56 GiB0.04 GiB9±22%
Phind-CodeLlama-34B-v2I1-IQ2_M33.7B10.72 GiB6.75 GiB18.56 GiB0.04 GiB9±22%
Skywork-R1V3-38BQ4_K_S38.4B17.49 GiB0.00 GiB18.56 GiB0.04 GiB9±22%
EXAONE-4.5-33BI1-Q3_K_M34.4B14.97 GiB2.49 GiB18.56 GiB0.04 GiB9±22%
Seed-OSS-36B-InstructUD-IQ1_M36.2B8.46 GiB9.00 GiB18.56 GiB0.04 GiB9±22%
MN-GRAND-23.5B-Gutenberg-UNCENSORED-V2-GLM4.7-ThinkingI1-IQ2_XXS23.4B6.12 GiB11.39 GiB18.55 GiB0.05 GiB9±22%
Qwen3-Coder-Next-REAMMoEI1-IQ2_XS60.3B16.71 GiB0.84 GiB18.55 GiB0.05 GiB40±37%
GRM-2.6-Plus-0628Q4_027.8B15.23 GiB2.25 GiB18.54 GiB0.06 GiB9±22%
ThinkingCap-Qwen3.6-27BQ4_027.4B15.23 GiB2.25 GiB18.54 GiB0.06 GiB9±22%
Tess-4-27BQ4_027.8B15.23 GiB2.25 GiB18.54 GiB0.06 GiB9±22%
Qwen3.8-27BIQ4_NL27.8B15.22 GiB2.25 GiB18.53 GiB0.07 GiB9±22%
Qwen3.6-27BIQ4_NL27.8B15.22 GiB2.25 GiB18.53 GiB0.07 GiB9±22%
Fimbulvetr-11B-v2Q8_010.7B10.74 GiB6.75 GiB18.52 GiB0.08 GiB9±22%
Huihui-Qwen3.5-35B-A3B-abliteratedMoEI1-Q3_K_L36.0B16.81 GiB0.70 GiB18.52 GiB0.08 GiB41±37%
Qwen3.5-35B-A3B-BaseMoEI1-Q3_K_L36.0B16.81 GiB0.70 GiB18.52 GiB0.08 GiB41±37%
Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-DistilledMoEI1-Q3_K_L36.0B16.81 GiB0.70 GiB18.52 GiB0.08 GiB41±37%
Qwen3-Coder-REAP-25B-A3BMoEQ4_K_M24.9B14.15 GiB3.38 GiB18.52 GiB0.08 GiB17±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation10.30 it/s7.6410.699
Benchmarked· n=9

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a RTX 4000 SFF Ada Generation run?
1793 of 2118 indexed open-weight models fit a RTX 4000 SFF Ada Generation at 131,072 context with q4_0 KV cache, the largest being Huihui-Qwen3-Coder-Next-abliterated at I1-IQ1_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a RTX 4000 SFF Ada Generation actually have?
Its nameplate is 20 GB, but about 18.60 GiB is available to a model once driver and compositor overhead is accounted for.
Is a RTX 4000 SFF Ada Generation fast for local AI?
Its memory bandwidth is 280 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.