NVIDIA · consumer

GeForce RTX 3080 Ti

GeForce RTX 3080 Ti has 20 GB of VRAM at 760 GB/s — about 18.60 GiB usable after driver and compositor overhead. 1955 of 2118 indexed models fit at 4K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
20 GB
GDDR6X
Bandwidth
760 GB/s
320-bit bus
Tensor FP16
136 TF
dense
TDP
350 W
$1199 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1678vision language 173video 16image 2audio asr 39audio tts 21embedding 26

What fits at 4K context

largest quantization that fits, per model · 1955 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Nemotron-Cascade-2-30B-A3BMoEQ3_K_M31.6B17.76 GiB0.06 GiB18.60 GiB0.00 GiB149±37%
OpenBuddy-R1-0528-Distill-Qwen3-32B-Preview0-QATQ4_032.8B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3-VL-32B-Instruct-ultra-uncensored-hereticI1-Q4_033.4B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
Huihui-Qwen3-VL-32B-Instruct-abliteratedI1-Q4_033.4B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
KAT-DevQ4_032.8B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
ColorGUI-32BI1-Q4_033.4B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3-VL-32B-InstructQ4_033.4B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3-VL-32B-ThinkingQ4_033.4B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3-32B-UncensoredI1-Q4_032.8B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3-32BQ4_032.8B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3-32B-abliteratedI1-Q4_032.8B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
DeepSWE-PreviewQ4_032.8B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
AReaL-boba-2-32BI1-Q4_032.8B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
Assistant_Pepe_32BI1-Q4_032.8B17.42 GiB0.28 GiB18.59 GiB0.01 GiB31±12.9%
granite-4.0-h-smallMoEQ4_K_S32.2B17.78 GiB0.02 GiB18.59 GiB0.01 GiB87±37%
t5-v1_1-xxlF324.8B17.74 GiB0.00 GiB18.59 GiB0.01 GiB31±12.9%
c4ai-command-r-08-2024Q4_032.3B17.49 GiB0.18 GiB18.58 GiB0.02 GiB31±12.9%
CallerIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
Dumpling-Qwen2.5-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
OREAL-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
INTELLECT-2IQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
openhands-lm-32b-v0.1IQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
LongWriter-Zero-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
OpenCodeReasoning-Nemotron-32B-IOIIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
OlympicCoder-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
OpenCodeReasoning-Nemotron-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
OpenThinker-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
QwQ-32B-ArliAI-RpR-v4IQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-Coder-32B-InstructIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-Coder-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
QwQ-32B-abliteratedIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
OpenThinker2-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
QwQ-32B-PreviewIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-32b-RP-InkIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
TinyR1-32B-PreviewIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
deepseek-r1-qwen-2.5-32B-ablatedIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
Rombos-LLM-V2.5-Qwen-32bIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
QwQ-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
DeepSeek-R1-Distill-Qwen-32B-abliteratedIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
DeepSeek-R1-Distill-Qwen-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-VL-32B-InstructIQ4_NL33.5B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
Delphi-25B-SimpleRL-MathI1-Q5_K_M25.0B16.52 GiB1.18 GiB18.58 GiB0.02 GiB31±12.9%
cogito-v1-preview-qwen-32BIQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
QwQ-32B-Snowdrop-v0IQ4_NL32.8B17.40 GiB0.28 GiB18.58 GiB0.02 GiB31±12.9%
Gemma-4-Dark-Gemistry-31BQ4_032.7B17.18 GiB0.51 GiB18.56 GiB0.04 GiB31±12.9%
Pantheon-Reasoning-27BQ4_K_L27.8B17.63 GiB0.07 GiB18.56 GiB0.04 GiB31±12.9%
Qwen3.5-27BQ4_K_L27.8B17.63 GiB0.07 GiB18.56 GiB0.04 GiB31±12.9%
GLM-4.7-Flash-REAP-23B-A3BMoEQ6_K23.0B17.69 GiB0.06 GiB18.55 GiB0.05 GiB116±37%
gemma-4-31B-itQ4_031.3B17.16 GiB0.51 GiB18.55 GiB0.05 GiB31±12.9%
Yi-34B-200K-DARE-megamerge-v8Q3_K_L34.4B17.40 GiB0.26 GiB18.55 GiB0.05 GiB31±12.9%
Nous-Hermes-2-Yi-34BI1-Q3_K_L34.4B17.40 GiB0.26 GiB18.55 GiB0.05 GiB31±12.9%
Nous-Capybara-limarpv3-34BI1-Q3_K_L34.4B17.40 GiB0.26 GiB18.55 GiB0.05 GiB31±12.9%
magnum-v2-32bQ4_K_S32.5B17.36 GiB0.28 GiB18.54 GiB0.06 GiB31±12.9%
GLM-4.7-FlashMoEQ4_131.2B17.67 GiB0.06 GiB18.54 GiB0.06 GiB133±37%
Qwen3-Coder-NextMoEUD-TQ1_079.7B17.64 GiB0.11 GiB18.54 GiB0.06 GiB185±37%
Salience-1.5-FlashMoEQ4_K_L31.1B17.63 GiB0.11 GiB18.53 GiB0.07 GiB131±37%
GLM-4.7-Flash-hereticMoEQ4_129.9B17.65 GiB0.06 GiB18.52 GiB0.08 GiB133±37%
Llama-3_3-Nemotron-Super-49B-v1_5IQ2_S49.9B14.76 GiB2.81 GiB18.51 GiB0.09 GiB31±12.9%
Valkyrie-49B-v2.1I1-IQ2_S49.9B14.76 GiB2.81 GiB18.51 GiB0.09 GiB31±12.9%
Llama-3_3-Nemotron-Super-49B-v1IQ2_S49.9B14.76 GiB2.81 GiB18.51 GiB0.09 GiB31±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a GeForce RTX 3080 Ti run?
1955 of 2118 indexed open-weight models fit a GeForce RTX 3080 Ti at 4,096 context with q4_0 KV cache, the largest being Nemotron-Cascade-2-30B-A3B at Q3_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3080 Ti actually have?
Its nameplate is 20 GB, but about 18.60 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3080 Ti fast for local AI?
Its memory bandwidth is 760 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.