NVIDIA · consumer

GeForce RTX 2080 Ti

GeForce RTX 2080 Ti has 11 GB of VRAM at 616 GB/s — about 10.23 GiB usable after driver and compositor overhead. 829 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
11 GB
GDDR6
Bandwidth
616 GB/s
352-bit bus
Tensor FP16
108 TF
dense
TDP
250 W
$999 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
audio asr 38text 643vision language 92audio tts 20embedding 21video 14image 1

What fits at 128K context

largest quantization that fits, per model · 829 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Voxtral-Mini-3B-2507IQ2_M4.7B1.45 GiB7.97 GiB10.23 GiB0.00 GiB46±12.9%
AMD-OLMo-1B-SFT-DPOQ6_K_L1.2B0.92 GiB8.50 GiB10.22 GiB0.01 GiB46±12.9%
Dolphin3.0-Llama3.2-3BQ4_K_L3.2B1.97 GiB7.44 GiB10.22 GiB0.01 GiB46±12.9%
Llama-Song-Stream-3B-InstructQ4_K_L3.2B1.97 GiB7.44 GiB10.22 GiB0.01 GiB46±12.9%
Llama-Doctor-3.2-3B-InstructQ4_K_L3.2B1.97 GiB7.44 GiB10.22 GiB0.01 GiB46±12.9%
llama-3.2-Korean-Bllossom-3BQ4_K_L3.2B1.97 GiB7.44 GiB10.22 GiB0.01 GiB46±12.9%
Llama-3.2-3B-InstructQ4_K_L3.2B1.97 GiB7.44 GiB10.22 GiB0.01 GiB46±12.9%
Hermes-3-Llama-3.2-3BQ4_K_L3.2B1.97 GiB7.44 GiB10.22 GiB0.01 GiB46±12.9%
granite-3.1-2b-instructQ6_K2.5B4.10 GiB5.31 GiB10.21 GiB0.02 GiB46±12.9%
Tiger-Gemma-12B-v3IQ3_XXS12.8B4.86 GiB4.50 GiB10.20 GiB0.03 GiB47±12.9%
AfriqueGemma-12BI1-IQ3_XXS12.2B4.86 GiB4.50 GiB10.20 GiB0.03 GiB47±12.9%
Llama-3.2-3B-Instruct-roleplay-tunedI1-Q4_13.2B1.95 GiB7.44 GiB10.20 GiB0.03 GiB47±12.9%
Llama-3.2-3B-Instruct-heretic-ablitered-uncensoredI1-Q4_13.2B1.95 GiB7.44 GiB10.20 GiB0.03 GiB47±12.9%
llama-3.2-3b-instruct-bnb-4bitQ4_13.3B1.95 GiB7.44 GiB10.20 GiB0.03 GiB47±12.9%
Llama3.2-3B-creative-writer-v0.1I1-Q4_13.2B1.95 GiB7.44 GiB10.20 GiB0.03 GiB47±12.9%
Firefly-V3.2I1-Q4_13.2B1.95 GiB7.44 GiB10.20 GiB0.03 GiB47±12.9%
Firefly-V3I1-Q4_13.2B1.95 GiB7.44 GiB10.20 GiB0.03 GiB47±12.9%
ARK-ASR-3BF164.1B6.99 GiB2.39 GiB10.19 GiB0.04 GiB47±12.9%
gemma-3-12b-it-vl-Gemini-3-Pro-Preview-Heretic-Uncensored-ThinkingI1-IQ3_XS12.2B4.85 GiB4.50 GiB10.19 GiB0.04 GiB47±12.9%
gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-ThinkingI1-IQ3_XS12.2B4.85 GiB4.50 GiB10.19 GiB0.04 GiB47±12.9%
gemma-3-12b-it-vl-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-IQ3_XS12.2B4.85 GiB4.50 GiB10.19 GiB0.04 GiB47±12.9%
gemma-3-12b-it-ultra-uncensored-hereticIQ3_XS12.2B4.85 GiB4.50 GiB10.19 GiB0.04 GiB47±12.9%
Floppa-12B-Gemma3-UncensoredI1-IQ3_XS12.2B4.85 GiB4.50 GiB10.19 GiB0.04 GiB47±12.9%
gemma-3-12b-it-hereticI1-IQ3_XS12.2B4.85 GiB4.50 GiB10.19 GiB0.04 GiB47±12.9%
gemma-3-12b-it-abliteratedIQ3_XS12.2B4.85 GiB4.50 GiB10.19 GiB0.04 GiB47±12.9%
Kepler-8B-Instruct-v2Q2_K7.6B5.61 GiB3.72 GiB10.19 GiB0.04 GiB47±12.9%
Darwin-4B-ChimeraF164.0B7.50 GiB1.87 GiB10.18 GiB0.05 GiB47±12.9%
Qwen3.6-28BMoEI1-IQ2_XS28.2B8.04 GiB1.33 GiB10.17 GiB0.06 GiB112±37%
Qwen3.5-28BMoEI1-IQ2_XS28.7B8.04 GiB1.33 GiB10.17 GiB0.06 GiB112±37%
orpheus-3b-0.1-pretrainedQ2_K_L3.8B1.92 GiB7.44 GiB10.16 GiB0.07 GiB47±12.9%
glm-4v-9bQ8_013.9B9.31 GiB0.00 GiB10.16 GiB0.07 GiB47±12.9%
Qwen3.6-35B-A3B-REAM-160-ru-agentMoEQ2_K_S23.6B8.02 GiB1.33 GiB10.16 GiB0.07 GiB109±37%
Llama-3.2-3B-Instruct-uncensoredIQ4_XS3.6B1.91 GiB7.44 GiB10.16 GiB0.07 GiB47±12.9%
Ling-mini-2.0MoEQ3_K_S16.3B6.71 GiB2.66 GiB10.15 GiB0.08 GiB75±37%
Gemma-4-12B-StyleTuneI1-Q2_K13.0B4.81 GiB4.50 GiB10.15 GiB0.08 GiB47±12.9%
gemma-4-12b-heretic-styletune-headI1-Q2_K12.0B4.81 GiB4.50 GiB10.15 GiB0.08 GiB47±12.9%
syrian-gemma-12bI1-Q2_K13.0B4.81 GiB4.50 GiB10.15 GiB0.08 GiB47±12.9%
internlm3-8b-instructQ5_K_L8.8B6.14 GiB3.19 GiB10.15 GiB0.08 GiB47±12.9%
Llama-3.2-3B-Instruct-abliteratedI1-IQ4_XS3.6B1.90 GiB7.44 GiB10.14 GiB0.09 GiB47±12.9%
Skywork-R1V3-38BIQ2_XS38.4B9.27 GiB0.00 GiB10.14 GiB0.09 GiB47±12.9%
Grug-12BIQ3_XXS12.0B4.79 GiB4.50 GiB10.14 GiB0.09 GiB47±12.9%
gemma-4-12B-it-Esper4IQ3_XXS12.0B4.79 GiB4.50 GiB10.14 GiB0.09 GiB47±12.9%
gemma-4-12B-itIQ3_XXS12.0B4.79 GiB4.50 GiB10.14 GiB0.09 GiB47±12.9%
Fara1.5-9BQ6_K9.4B7.17 GiB2.13 GiB10.13 GiB0.10 GiB47±12.9%
QwenPaw-Flash-9BQ6_K9.4B7.17 GiB2.13 GiB10.13 GiB0.10 GiB47±12.9%
grug-9bQ6_K9.4B7.17 GiB2.13 GiB10.13 GiB0.10 GiB47±12.9%
OmniCoder-9BQ6_K9.4B7.17 GiB2.13 GiB10.13 GiB0.10 GiB47±12.9%
Ornith-1.0-9BQ6_K9.2B7.17 GiB2.13 GiB10.13 GiB0.10 GiB47±12.9%
Qwen3.5-9B-NeoQ6_K9.7B7.17 GiB2.13 GiB10.13 GiB0.10 GiB47±12.9%
gemma-4-12B-it-hereticQ3_K_S12.0B4.78 GiB4.50 GiB10.13 GiB0.10 GiB47±12.9%
Llama-3.2-3B-Instruct-uncensoredQ4_K_M3.2B1.88 GiB7.44 GiB10.13 GiB0.10 GiB47±12.9%
Llama-3.2-3BQ4_K_M3.2B1.88 GiB7.44 GiB10.13 GiB0.10 GiB47±12.9%
OneLLM-Doey-ChatQA-V1-Llama-3.2-3BQ4_K_M3.2B1.88 GiB7.44 GiB10.13 GiB0.10 GiB47±12.9%
Llama-3.2-3B-bnb-4bitQ4_K_M3.3B1.88 GiB7.44 GiB10.13 GiB0.10 GiB47±12.9%
Qwen3.5-9BQ6_K9.7B7.16 GiB2.13 GiB10.12 GiB0.11 GiB47±12.9%
Ministral-3-3B-Instruct-2512-BF16Q5_K_L4.3B2.39 GiB6.91 GiB10.11 GiB0.12 GiB47±12.9%
orpheus-3b-0.1-ftQ4_K_S3.8B1.86 GiB7.44 GiB10.11 GiB0.12 GiB47±12.9%
Muse-Glimmer-30BIQ2_XXS29.8B8.31 GiB0.91 GiB10.10 GiB0.13 GiB47±12.9%
nomic-embed-codeQ6_K_L7.1B5.53 GiB3.72 GiB10.10 GiB0.13 GiB47±12.9%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopQ5_012.1B7.65 GiB1.59 GiB10.10 GiB0.13 GiB47±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation11.62 it/s8.9613.801,506
Benchmarked· n=1,506

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 2080 Ti run?
829 of 2118 indexed open-weight models fit a GeForce RTX 2080 Ti at 131,072 context with q8_0 KV cache, the largest being Voxtral-Mini-3B-2507 at IQ2_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 2080 Ti actually have?
Its nameplate is 11 GB, but about 10.23 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 2080 Ti fast for local AI?
Its memory bandwidth is 616 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.