NVIDIA · datacenter

H200 SXM

H200 SXM has 141 GB of VRAM at 4800 GB/s — about 131.13 GiB usable after driver and compositor overhead. 2088 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
141 GB
HBM3e
Bandwidth
4800 GB/s
6144-bit bus
Tensor FP16
989 TF
dense
TDP
700 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1793vision language 191image 2audio tts 21audio asr 39video 16embedding 26

What fits at 128K context

largest quantization that fits, per model · 2088 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
step-3.5-flashIQ4_NL199B103.97 GiB26.05 GiB131.05 GiB0.08 GiB21±22%
MiniMax-M2.5MoEIQ4_XS229B113.53 GiB16.47 GiB130.98 GiB0.15 GiB57±37%
MiniMax-M2.1MoEI1-IQ4_XS229B113.51 GiB16.47 GiB130.96 GiB0.17 GiB57±37%
Qwen3-235B-A22B-Instruct-2507MoEIQ4_XS235B117.24 GiB12.48 GiB130.75 GiB0.38 GiB58±37%
Qwen3-235B-A22B-Thinking-2507MoEIQ4_XS235B117.24 GiB12.48 GiB130.75 GiB0.38 GiB58±37%
Qwen3-235B-A22BMoEIQ4_XS235B116.89 GiB12.48 GiB130.40 GiB0.73 GiB58±37%
MiMo-V2.5MoEKV unresolvedIQ3_XXS311B121.31 GiB7.97 GiB130.32 GiB0.81 GiB80±37%
Qwen3-VL-235B-A22B-ThinkingMoEIQ4_XS236B116.70 GiB12.48 GiB130.21 GiB0.92 GiB58±37%
Qwen3-VL-235B-A22B-InstructMoEIQ4_XS236B116.70 GiB12.48 GiB130.21 GiB0.92 GiB58±37%
DeepSeek-Coder-V2-Instruct-0724MoEQ4_K_S236B124.68 GiB4.48 GiB130.20 GiB0.93 GiB91±37%
DeepSeek-V2.5MoEQ4_K_S236B124.68 GiB4.48 GiB130.20 GiB0.93 GiB91±37%
DeepSeek-Coder-V2-InstructMoEQ4_K_S236B124.68 GiB4.48 GiB130.20 GiB0.93 GiB91±37%
Qwen3-235B-A22B-abliteratedMoEI1-IQ4_XS235B116.68 GiB12.48 GiB130.20 GiB0.93 GiB58±37%
DeepSeek-V3-0324MoEIQ1_S685B124.38 GiB4.56 GiB130.02 GiB1.11 GiB95±37%
DeepSeek-R1MoEIQ1_S685B124.38 GiB4.56 GiB130.02 GiB1.11 GiB95±37%
MiniMax-M3MoEIQ2_XS427B120.63 GiB7.97 GiB129.61 GiB1.52 GiB80±37%
WizardLM-Uncensored-SuperCOT-StoryTelling-30bQ6_K32.5B24.85 GiB103.59 GiB129.52 GiB1.61 GiB21±22%
Wizard-Vicuna-30B-UncensoredI1-Q6_K32.5B24.85 GiB103.59 GiB129.52 GiB1.61 GiB21±22%
archangel_sft-kto_llama30bI1-Q6_K32.5B24.85 GiB103.59 GiB129.52 GiB1.61 GiB21±22%
DeepSeek-V4-FlashMoEUD-IQ4_NL291B128.43 GiB0.03 GiB129.51 GiB1.62 GiB137±37%
command-a-plus-05-2026-bf16MoEQ4_K_L219B126.06 GiB2.35 GiB129.41 GiB1.72 GiB86±37%
Trinity-Large-ThinkingMoEIQ2_M399B123.88 GiB4.40 GiB129.31 GiB1.82 GiB108±37%
dots.llm1.instMoEIQ3_XS143B61.71 GiB65.88 GiB128.61 GiB2.52 GiB22±37%
DeepSeek-V4-Flash-0731MoEUD-IQ4_NL304B127.28 GiB0.03 GiB128.36 GiB2.77 GiB138±37%
Trinity-Large-TrueBaseMoEI1-Q2_K_S399B122.88 GiB4.40 GiB128.31 GiB2.82 GiB109±37%
GLM-4.6-REAP-268B-A32BMoEUD-IQ3_XXS269B102.71 GiB24.44 GiB128.18 GiB2.95 GiB42±37%
grok-2MoEQ3_K_S270B109.94 GiB17.00 GiB128.08 GiB3.05 GiB31±37%
MiMo-V2-FlashMoEKV unresolvedUD-IQ3_XXS310B118.70 GiB7.97 GiB127.72 GiB3.41 GiB81±37%
MiniMax-M2.7-BF16-ultra-uncensored-hereticMoEQ3_K_L229B110.22 GiB16.47 GiB127.68 GiB3.45 GiB58±37%
Ornith-1.0-397BMoEIQ2_M397B123.99 GiB1.99 GiB127.03 GiB4.10 GiB125±37%
Step-3.7-FlashIQ4_XS201B99.94 GiB26.05 GiB127.01 GiB4.12 GiB22±22%
granite-34b-code-base-8kF3233.7B125.60 GiB0.00 GiB126.68 GiB4.45 GiB22±22%
NVIDIA-Nemotron-3-Super-120B-A12B-BF16MoEQ8_0124B119.65 GiB5.84 GiB126.49 GiB4.64 GiB84±37%
Llama-4-Maverick-17B-128E-InstructMoEKV unresolvedUD-IQ1_S402B112.48 GiB12.75 GiB126.25 GiB4.88 GiB74±37%
Qwen3.5-122B-A10BMoEQ8_0125B123.49 GiB1.59 GiB126.11 GiB5.02 GiB116±37%
GLM-4.5MoEUD-IQ1_M358B100.32 GiB24.44 GiB125.80 GiB5.33 GiB44±37%
GLM-4.7MoEUD-IQ1_M358B100.27 GiB24.44 GiB125.74 GiB5.39 GiB44±37%
Mistral-Medium-3.5-128BQ6_K_L128B101.13 GiB23.38 GiB125.66 GiB5.47 GiB22±22%
GLM-4.6MoEUD-IQ1_M357B100.02 GiB24.44 GiB125.50 GiB5.63 GiB44±37%
Solar-Open2-250BMoEQ3_K_M250B111.63 GiB12.75 GiB125.41 GiB5.72 GiB69±37%
Qwen3-Coder-REAP-363B-A35BMoEUD-IQ1_M363B107.85 GiB16.47 GiB125.35 GiB5.78 GiB52±37%
Laguna-S-2.1MoEQ8_0118B119.91 GiB3.26 GiB124.19 GiB6.94 GiB100±37%
MiniMax-M2.1-REAP-139B-A10BMoEI1-Q6_K139B106.40 GiB16.47 GiB123.86 GiB7.27 GiB54±37%
m51Lab-MiniMax-M2.7-REAP-139B-A10BMoEQ6_K139B106.40 GiB16.47 GiB123.86 GiB7.27 GiB54±37%
HunyuanImage-2.1Q8_017.5B122.73 GiB0.00 GiB123.78 GiB7.35 GiB22±22%
Qwen3.5-122B-A10B-hereticMoEQ8_0123B120.95 GiB1.59 GiB123.57 GiB7.56 GiB118±37%
Hy3MoEQ2_K299B101.28 GiB21.25 GiB123.57 GiB7.56 GiB49±37%
Mixtral-8x22B-v0.1MoEQ6_K141B107.60 GiB14.88 GiB123.53 GiB7.60 GiB33±37%
GLM-4.7-REAP-218B-A32BMoEQ3_K_M218B97.57 GiB24.44 GiB123.05 GiB8.08 GiB41±37%
Hermes-4-405BUD-IQ1_M406B88.23 GiB33.47 GiB122.98 GiB8.15 GiB22±22%
GLM-4.5-AirMoEQ8_0110B109.39 GiB12.22 GiB122.64 GiB8.49 GiB60±37%
GLM-4.5-Air-DerestrictedMoEQ8_0110B109.39 GiB12.22 GiB122.64 GiB8.49 GiB60±37%
Qwen3.5-REAP-212B-A17BMoEQ4_K_M212B119.56 GiB1.99 GiB122.60 GiB8.53 GiB110±37%
Trinity-Large-PreviewMoEIQ2_M399B116.50 GiB4.40 GiB121.93 GiB9.20 GiB112±37%
Hermes-3-Llama-3.1-405BIQ1_M406B87.08 GiB33.47 GiB121.83 GiB9.30 GiB23±22%
Qwen3.5-397B-A17BMoEIQ2_S403B118.57 GiB1.99 GiB121.61 GiB9.52 GiB129±37%
c4ai-command-r-plus-08-2024Q8_0104B102.74 GiB17.00 GiB120.92 GiB10.21 GiB23±22%
command-r-35b-writer-v2Q8_035.0B34.63 GiB85.00 GiB120.73 GiB10.40 GiB23±22%
MiniMax-M2.7MoEUD-IQ4_NL229B103.15 GiB16.47 GiB120.60 GiB10.53 GiB59±37%
Llama-4-Scout-17B-16E-InstructMoEKV unresolvedQ8_0109B106.67 GiB12.75 GiB120.44 GiB10.69 GiB60±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a H200 SXM run?
2088 of 2118 indexed open-weight models fit a H200 SXM at 131,072 context with q8_0 KV cache, the largest being step-3.5-flash at IQ4_NL. That covers text, vision-language, image, video and speech models.
How much usable memory does a H200 SXM actually have?
Its nameplate is 141 GB, but about 131.13 GiB is available to a model once driver and compositor overhead is accounted for.
Is a H200 SXM fast for local AI?
Its memory bandwidth is 4800 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.