NVIDIA · consumer

GeForce RTX 3080 Ti

GeForce RTX 3080 Ti has 20 GB of VRAM at 760 GB/s — about 18.60 GiB usable after driver and compositor overhead. 1798 of 2118 indexed models fit at 64K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
20 GB
GDDR6X
Bandwidth
760 GB/s
320-bit bus
Tensor FP16
136 TF
dense
TDP
350 W
$1199 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1524vision language 170audio tts 21image 2video 16audio asr 39embedding 26

What fits at 64K context

largest quantization that fits, per model · 1798 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
GLM-4.7-FlashMoEQ4_031.2B16.03 GiB1.76 GiB18.60 GiB0.00 GiB82±37%
GLM-4.7-Flash-hereticMoEIQ4_NL29.9B16.03 GiB1.76 GiB18.60 GiB0.00 GiB82±37%
Gemma4-Gutenberg-31BIQ2_M31.3B11.78 GiB5.94 GiB18.60 GiB0.00 GiB31±12.9%
gemma-4-31B-itIQ2_M31.3B11.78 GiB5.94 GiB18.60 GiB0.00 GiB31±12.9%
Gemma4-Gutenberg-31B-HereticIQ2_M31.3B11.78 GiB5.94 GiB18.60 GiB0.00 GiB31±12.9%
Equinox-31BIQ2_M31.3B11.78 GiB5.94 GiB18.60 GiB0.00 GiB31±12.9%
gemma-4-31B-it-SDFT-Heretic-RPIQ2_M30.7B11.78 GiB5.94 GiB18.60 GiB0.00 GiB31±12.9%
dolphin-2.9.3-mistral-7B-32kF167.2B13.50 GiB4.25 GiB18.59 GiB0.01 GiB31±12.9%
Mistral-7B-Instruct-v0.3-ParasiteF167.2B13.50 GiB4.25 GiB18.59 GiB0.01 GiB31±12.9%
Mistral-7B-Instruct-v0.3-JbliteratedF167.2B13.50 GiB4.25 GiB18.59 GiB0.01 GiB31±12.9%
Mistral-7B-Instruct-v0.3F167.2B13.50 GiB4.25 GiB18.59 GiB0.01 GiB31±12.9%
Mistral-7B-v0.3F167.2B13.50 GiB4.25 GiB18.59 GiB0.01 GiB31±12.9%
Mistral-7B-v0.3-Chinese-ChatF167.2B13.50 GiB4.25 GiB18.59 GiB0.01 GiB31±12.9%
mistral-7b-v0.3-bnb-4bitBF167.5B13.50 GiB4.25 GiB18.59 GiB0.01 GiB31±12.9%
Mathstral-7B-v0.1F167.2B13.50 GiB4.25 GiB18.59 GiB0.01 GiB31±12.9%
Tess-4-9BF169.7B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Crow-9B-HERETIC-4.6F169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-Claude-4.6-HighIQ-INSTRUCT-HERETIC-UNCENSOREDF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSOREDF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-Claude-4.6-OS-HERETIC-UNCENSORED-INSTRUCTF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwable-9B-Claude-Fable-5-hereticF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Holo-3.1-9BF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwable-9B-Claude-Fable-5F169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwythos-9B-Claude-Mythos-5-1M-uncensored-hereticBF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-ultra-uncensored-hereticF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-abliterated-v2-MAXF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwable-9B-Claude-Fable-5-StraTAF169.0B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-Uncensored-cyber-v3F169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
NaNovel-9BF169.7B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
OmniCoder-9B-Claude-Opus-High-Reasoning-DistillF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-Fable-5-v1BF169.7B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwable-9B-Claude-Fable-5-OBLITERATEDF169.0B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-Unredacted-MAXF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Huihui-Qwen3.5-9B-abliteratedF169.7B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-abliteratedF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
PlutoF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
QwenPaw-Flash-9BF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Holo-3.1-9B-CoderF169.0B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Holo-3.1-9B-abliterated-rdoF169.0B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
MaralGPT-Mythos-9B-2606BF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
qwen3.5-9b-nsfw-captioning-v5F169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Fara1.5-9BBF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
grug-9bBF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
OmniCoder-9BBF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-BaseF169.7B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-DS-v4-Flash-v3.0BF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKINGBF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-Star-Trek-TNG-DS9-Heretic-Uncensored-ThinkingBF169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-heretic-v2F169.4B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-Claude-Opus-4.6-DistillF169.0B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Miss_MARTHA-9B-Qwen3.5-OmniF169.0B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-DeepSeek-V4-FlashF169.7B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Huihui-Qwen3.5-9B-Claude-4.6-Opus-abliteratedF169.7B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.5-9B-NeoBF169.7B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Qwopus3.5-9B-v3.5BF169.7B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
Katarau-9B-ru-RP-nsfwF169.0B16.69 GiB1.06 GiB18.59 GiB0.01 GiB31±12.9%
t5-v1_1-xxlF324.8B17.74 GiB0.00 GiB18.59 GiB0.01 GiB31±12.9%
Qwen3.6-27B-Fable-5-ExperimentalIQ4_NL27.8B15.59 GiB2.13 GiB18.58 GiB0.02 GiB31±12.9%
OpenChat-3.5-7B-Qwen-v2.0KV unresolvedF167.2B13.49 GiB4.25 GiB18.58 GiB0.02 GiB31±12.9%
ContextualKunoichi_KTO-7BF167.2B13.49 GiB4.25 GiB18.58 GiB0.02 GiB31±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a GeForce RTX 3080 Ti run?
1798 of 2118 indexed open-weight models fit a GeForce RTX 3080 Ti at 65,536 context with q8_0 KV cache, the largest being GLM-4.7-Flash at Q4_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3080 Ti actually have?
Its nameplate is 20 GB, but about 18.60 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3080 Ti fast for local AI?
Its memory bandwidth is 760 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.