NVIDIA · datacenter

A100 40GB

A100 40GB has 40 GB of VRAM at 1555 GB/s — about 37.20 GiB usable after driver and compositor overhead. 2003 of 2118 indexed models fit at 64K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
40 GB
HBM2
Bandwidth
1555 GB/s
5120-bit bus
Tensor FP16
312 TF
dense
TDP
400 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 180text 1719video 16audio tts 21audio asr 39image 2embedding 26

What fits at 64K context

largest quantization that fits, per model · 2003 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3.6-27B-uncensored-heretic-v2Q8_027.4B33.96 GiB2.13 GiB37.14 GiB0.06 GiB25±22%
Hunyuan-A13B-InstructMoEUD-IQ3_XXS80.4B31.88 GiB4.25 GiB37.13 GiB0.07 GiB24±22%
deepseek-llm-67b-chatI1-Q2_K67.4B23.40 GiB12.62 GiB37.11 GiB0.09 GiB25±22%
deepseek-llm-67b-baseI1-Q2_K67.4B23.40 GiB12.62 GiB37.11 GiB0.09 GiB25±22%
openbuddy-deepseek-67b-v15.3-4kI1-Q2_K67.4B23.40 GiB12.62 GiB37.11 GiB0.09 GiB25±22%
Yi-1.5-9B-ChatF328.8B32.89 GiB3.19 GiB37.11 GiB0.09 GiB25±22%
Qwen3-Coder-Next-REAMMoEI1-Q4_160.3B35.30 GiB0.80 GiB37.09 GiB0.11 GiB125±37%
Step-3.5-Flash-REAP-121B-A11BI1-IQ1_S121B22.74 GiB13.30 GiB37.06 GiB0.14 GiB25±22%
Behemoth-X-123B-v2IQ1_S123B24.18 GiB11.69 GiB37.02 GiB0.18 GiB25±22%
Llama-4-Scout-17B-16E-Instruct-abliterated-v2MoEKV unresolvedI1-IQ2_XS109B29.60 GiB6.38 GiB37.01 GiB0.19 GiB50±37%
OLMo-2-1124-13B-InstructQ5_K_L13.7B9.39 GiB26.56 GiB36.99 GiB0.21 GiB25±22%
Ornith-1.0-35B-AEON-Ultimate-Uncensored-BF16MoEIQ4_XS35.1B35.30 GiB0.66 GiB36.97 GiB0.23 GiB124±37%
Rombo-LLM-V3.0-Qwen-72bI1-IQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Qwen2.5-72B-Instruct-abliteratedI1-IQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Qwen2.5-72B-Instruct-abliterated-v2I1-IQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
HuatuoGPT-o1-72BIQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
MiroThinker-v1.0-72BI1-IQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
EVA-Qwen2.5-72B-v0.2IQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Qwen2.5-Math-72B-InstructIQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Qwen2.5-72B-InstructIQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Malaysian-Qwen2.5-72B-InstructI1-IQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Qwen2.5-72BI1-IQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
magnum-v4-72bI1-IQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
KAT-Dev-72B-ExpIQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Homer-v1.0-Qwen2.5-72BIQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Tower-Plus-72B-ultra-uncensored-hereticI1-IQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Qwen2.5-VL-72B-InstructIQ2_XS73.4B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Chronos-Platinum-72BIQ2_XS72.7B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
UI-TARS-72B-DPOIQ2_XS73.4B25.20 GiB10.63 GiB36.95 GiB0.25 GiB25±22%
Llama3.2-30B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEQ8_030.0B29.66 GiB6.24 GiB36.91 GiB0.29 GiB43±37%
Salience-1.5-ProMoEQ8_036.0B35.22 GiB0.66 GiB36.89 GiB0.31 GiB124±37%
Qwable-v1MoEQ8_036.0B35.22 GiB0.66 GiB36.89 GiB0.31 GiB124±37%
T-SearchMoEQ8_036.0B35.22 GiB0.66 GiB36.89 GiB0.31 GiB124±37%
Qwen3.5-35B-A3BMoEQ8_036.0B35.22 GiB0.66 GiB36.88 GiB0.32 GiB124±37%
Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-PreservedMoEQ8_035.1B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Qwen3.6-35B-A3B-Fable-5-DistillMoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Qwable-v2MoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Qwen3.6-35B-A3B-YOYO-V2MoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Ornith-1.0-35B-FP8-BLOCK-MTPMoEQ8_035.5B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
fable-coder-35B-A3BMoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
PINQWEN-3.6-35B-CLEAN-BF16MoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
UniMath-35B-A3BMoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Ornith-1.0-35B-Heretic-MTPMoEQ8_035.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Fawen-1.0-35BMoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Qwopus3.6-35B-A3B-v1MoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
CyberStrike-OffSec-35BMoEQ8_035.1B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Qwen3.6-35B-A3BMoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-DistilledMoEQ8_036.0B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Qwen3.5-35B-A3B-uncensored-heretic-v2-Native-MTP-PreservedMoEQ8_035.1B35.21 GiB0.66 GiB36.87 GiB0.33 GiB124±37%
Huihui-GLM-4.7-Flash-abliterated-57BMoEI1-Q4_K_M57.3B31.37 GiB4.45 GiB36.86 GiB0.34 GiB60±37%
Assistant_Pepe_70BIQ2_M70.6B25.10 GiB10.63 GiB36.85 GiB0.35 GiB25±22%
Qwen3.6-34B-80L-Fable-5-HereticQ8_033.4B33.08 GiB2.66 GiB36.80 GiB0.40 GiB25±22%
14BQ5_014.2B9.17 GiB26.56 GiB36.78 GiB0.42 GiB25±22%
NSFW_13B_sftQ5_K_M13.3B9.17 GiB26.56 GiB36.78 GiB0.42 GiB25±22%
Mistral-Small-4-119B-2603MoEUD-IQ2_M119B34.99 GiB0.75 GiB36.77 GiB0.43 GiB122±37%
Qwen2.5-Coder-14B-InstructQ8_014.8B29.25 GiB6.38 GiB36.67 GiB0.53 GiB25±22%
Qwen3-Coder-NextMoEQ3_K_S79.7B32.47 GiB3.19 GiB36.65 GiB0.55 GiB85±37%
Qwen3-Next-80B-A3B-ThinkingMoEQ3_K_S81.3B32.47 GiB3.19 GiB36.65 GiB0.55 GiB85±37%
Qwen3-Next-80B-A3B-InstructMoEQ3_K_S81.3B32.47 GiB3.19 GiB36.65 GiB0.55 GiB85±37%
Noromaid-v0.4-Mixtral-Instruct-8x7b-ZlossMoEQ5_K_M46.7B31.34 GiB4.25 GiB36.62 GiB0.58 GiB37±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a A100 40GB run?
2003 of 2118 indexed open-weight models fit a A100 40GB at 65,536 context with q8_0 KV cache, the largest being Qwen3.6-27B-uncensored-heretic-v2 at Q8_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a A100 40GB actually have?
Its nameplate is 40 GB, but about 37.20 GiB is available to a model once driver and compositor overhead is accounted for.
Is a A100 40GB fast for local AI?
Its memory bandwidth is 1555 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.