NVIDIA · consumer

GeForce RTX 3080 Ti

GeForce RTX 3080 Ti has 20 GB of VRAM at 760 GB/s — about 18.60 GiB usable after driver and compositor overhead. 1603 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
20 GB
GDDR6X
Bandwidth
760 GB/s
320-bit bus
Tensor FP16
136 TF
dense
TDP
350 W
$1199 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1343vision language 157audio asr 39video 16embedding 26image 1audio tts 21

What fits at 128K context

largest quantization that fits, per model · 1603 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
ultravox-v0_5-llama-3_2-1bQ4_K_M683M0.75 GiB17.00 GiB18.60 GiB0.00 GiB31±12.9%
CycleGRPO-4BF164.8B8.23 GiB9.56 GiB18.60 GiB0.00 GiB31±12.9%
EXAONE-4.0-32BQ3_K_S32.0B13.00 GiB4.70 GiB18.60 GiB0.00 GiB31±12.9%
Jan-v3-4B-base-instructBF164.4B8.22 GiB9.56 GiB18.60 GiB0.00 GiB31±12.9%
Jan-code-4bBF164.4B8.22 GiB9.56 GiB18.60 GiB0.00 GiB31±12.9%
dolphin-2.9.2-Phi-3-MediumKV unresolvedIQ2_M14.0B4.45 GiB13.28 GiB18.59 GiB0.01 GiB31±12.9%
t5-v1_1-xxlF324.8B17.74 GiB0.00 GiB18.59 GiB0.01 GiB31±12.9%
EVA-abliterated-TIES-Qwen2.5-14BI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Neuron-V1-14B-InstructI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Ektome-Qwen2.5-Coder-14B-Instruct-PristinelyUncensoredI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-14B-Instruct-1M-abliteratedI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
DeepCoder-14B-PreviewIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Deepseeker-Kunou-Qwen2.5-14bI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
SuperNova-MediusIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
14B-Qwen2.5-Kunou-v1I1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Sugoi-14B-Ultra-HFI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-Coder-14B-Instruct-abliteratedIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
OpenCodeReasoning-Nemotron-14BIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-14B-InstructIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
DeepSeek-R1-Distill-Qwen-14B-abliterated-v2I1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
C1-TachuI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
DeepSeek-R1-Distill-Qwen-14B-abliteratedI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
0x-liteIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-Coder-14B-InstructIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Tessera-4I1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-14B-InstructIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Tessera-4.1I1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-14B-Instruct-1MIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Qwen2.5-Coder-14BIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
AceReason-Nemotron-14BI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
DeepSeek-R1-Distill-Qwen-14BIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
UwU-14B-Math-v0.2I1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
EVA-Qwen2.5-14B-v0.2I1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
EVA-Qwen2.5-14B-v0.0I1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
EVA-Qwen2.5-14B-v0.1I1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
oxy-1-smallIQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Impish_QWEN_14B-1MI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Phi-3-medium-128k-instructQ2_K_S14.0B4.44 GiB13.28 GiB18.58 GiB0.02 GiB31±12.9%
Lamarck-14B-v0.7I1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
QwenStock-14BI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
DeepSeek-R1-Distill-Qwen-14B-UncensoredI1-IQ2_M14.8B4.99 GiB12.75 GiB18.58 GiB0.02 GiB31±12.9%
Smilodon-9B-v1I1-Q5_K_M10.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
bella-bartender-v2I1-Q5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-Q5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-Q5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
Gemma-2-9B-It-SPPO-Iter3I1-Q5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
Gemma-SEA-LION-v3-9B-ITI1-Q5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-Q5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
G2-Darkest-Writer-9B-v1I1-Q5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
Tiger-Gemma-9B-v3I1-Q5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
gemma-2-9b-it-abliteratedQ5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
gemma-2-9b-itQ5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
Tiger-Gemma-9B-v1Q5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
magnum-v4-9bQ5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
gemma-2-9bQ5_K_M9.2B6.19 GiB11.55 GiB18.58 GiB0.02 GiB31±12.9%
Llama3.2-24B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEI1-Q4_K_M18.0B10.33 GiB7.44 GiB18.58 GiB0.02 GiB33±37%
HomunculusQ4_K_M12.5B7.10 GiB10.63 GiB18.57 GiB0.03 GiB31±12.9%
Fimbulvetr-11B-v2I1-Q3_K_M10.7B4.98 GiB12.75 GiB18.57 GiB0.03 GiB31±12.9%
ERNIE-21B-A3B-Thinking-Gemini-3-Pro-High-Reasoning-V2I1-Q5_K_S21.8B14.02 GiB3.72 GiB18.56 GiB0.04 GiB31±12.9%
ERNIE-21B-A3B-Claude-4.5-High-OPUS-ThinkingI1-Q5_K_S21.8B14.02 GiB3.72 GiB18.56 GiB0.04 GiB31±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a GeForce RTX 3080 Ti run?
1603 of 2118 indexed open-weight models fit a GeForce RTX 3080 Ti at 131,072 context with q8_0 KV cache, the largest being ultravox-v0_5-llama-3_2-1b at Q4_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3080 Ti actually have?
Its nameplate is 20 GB, but about 18.60 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3080 Ti fast for local AI?
Its memory bandwidth is 760 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.