NVIDIA · consumer

GeForce RTX 2080 Ti

GeForce RTX 2080 Ti has 11 GB of VRAM at 616 GB/s — about 10.23 GiB usable after driver and compositor overhead. 1679 of 2118 indexed models fit at 16K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
11 GB
GDDR6
Bandwidth
616 GB/s
352-bit bus
Tensor FP16
108 TF
dense
TDP
250 W
$999 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1445vision language 132image 2video 14audio tts 21audio asr 39embedding 26

What fits at 16K context

largest quantization that fits, per model · 1679 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Jan-v3-4B-base-instructBF164.4B8.22 GiB1.20 GiB10.23 GiB0.00 GiB46±12.9%
Jan-code-4bBF164.4B8.22 GiB1.20 GiB10.23 GiB0.00 GiB46±12.9%
Rocinante-XL-16B-v1Q3_K_M16.1B7.59 GiB1.79 GiB10.23 GiB0.00 GiB47±12.9%
ERNIE-21B-A3B-Thinking-Gemini-3-Pro-High-Reasoning-V2I1-IQ3_M21.8B8.94 GiB0.46 GiB10.23 GiB0.00 GiB46±12.9%
ERNIE-21B-A3B-Claude-4.5-High-OPUS-ThinkingI1-IQ3_M21.8B8.94 GiB0.46 GiB10.23 GiB0.00 GiB46±12.9%
ERNIE-4.5-21B-A3B-ThinkingI1-IQ3_M21.8B8.94 GiB0.46 GiB10.23 GiB0.00 GiB46±12.9%
deepseek-coder-6.7b-instructQ6_K6.7B5.15 GiB4.25 GiB10.23 GiB0.00 GiB46±12.9%
deepseek-coder-6.7b-baseQ6_K6.7B5.15 GiB4.25 GiB10.23 GiB0.00 GiB46±12.9%
deepseek-coder-6.7B-kexerI1-Q6_K6.7B5.15 GiB4.25 GiB10.23 GiB0.00 GiB46±12.9%
Magicoder-S-DS-6.7BI1-Q6_K6.7B5.15 GiB4.25 GiB10.23 GiB0.00 GiB46±12.9%
Apriel-1.6-15b-ThinkerI1-Q4_K_S14.9B7.78 GiB1.59 GiB10.22 GiB0.01 GiB47±12.9%
MathCoder2-CodeLlama-7BQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
CodeLlama-7b-instruct-hfQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
CodeLlama-7b-hfQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
WizardLM-7B-UncensoredI1-Q6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
Llama-2-7B-32K-InstructI1-Q6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
Luna-AI-Llama2-UncensoredI1-Q6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
Llama-2-7b-chat-hfQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
Swallow-7b-NVE-instruct-hfI1-Q6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
llava-v1.5-7bQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
CodeLlama-7b-python-hfQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
Wizard-Vicuna-7B-UncensoredQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
llama2_7b_chat_uncensoredQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
WizardLM-7B-V1.0-UncensoredQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
Llama-2-7b-hfQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
pygmalion-2-7bQ6_K6.7B5.15 GiB4.25 GiB10.22 GiB0.01 GiB46±12.9%
Qwen3-VL-32B-InstructUD-IQ1_S33.4B7.21 GiB2.13 GiB10.22 GiB0.01 GiB47±12.9%
Qwen3-VL-32B-ThinkingUD-IQ1_S33.4B7.21 GiB2.13 GiB10.22 GiB0.01 GiB47±12.9%
Qwen3-32BUD-IQ1_S32.8B7.21 GiB2.13 GiB10.22 GiB0.01 GiB47±12.9%
Crow-9B-HERETIC-4.6Q8_09.4B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Qwen3.5-9B-CoderQ8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Qwopus3.5-9B-v3.5Q8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Qwythos-9B-Claude-Mythos-5-1M-MTPQ8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliteratedQ8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Qwen3.5-9B-Fable-5-v1Q8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
PINQWEN-3.5-9B-1M-BF16Q8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Openprose-2-FlashQ8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Qwen3.5-9B-Nikusui-v1Q8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Qwen3.5-9BQ8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Ornith-1.0-9B-heretic-MTPQ8_09.4B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
dotwebs-1Q8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
liftQ8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Ornith-1.0-9BQ8_09.2B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Qwen3.5-9B-DeepSeek-V4-FlashQ8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Qwen3.5-9BQ8_09.7B9.11 GiB0.27 GiB10.22 GiB0.01 GiB47±12.9%
Snowpiercer-15B-v4-hereticIQ4_XS15.0B7.70 GiB1.66 GiB10.21 GiB0.02 GiB47±12.9%
GLM-4.7-Flash-hereticMoEIQ2_S29.9B8.96 GiB0.44 GiB10.21 GiB0.02 GiB152±37%
Qwen3-Coder-REAP-25B-A3BMoEQ2_K_L24.9B8.61 GiB0.80 GiB10.20 GiB0.03 GiB121±37%
L3-DARKEST-PLANET-16.5BQ3_K_S16.5B7.00 GiB2.36 GiB10.20 GiB0.03 GiB47±12.9%
Qwen3-14B-Claude-4.5-Opus-High-Reasoning-DistillIQ4_NL14.8B8.01 GiB1.33 GiB10.19 GiB0.04 GiB47±12.9%
DeepSeek-R1-Distill-Llama-8B-AbliteratedI1-IQ4_XS8.0B8.28 GiB1.06 GiB10.19 GiB0.04 GiB47±12.9%
Orca-2-13b-Alpaca-UncensoredI1-IQ1_S13.0B2.70 GiB6.64 GiB10.18 GiB0.05 GiB47±12.9%
WizardLM-13B-UncensoredI1-IQ1_S13.0B2.70 GiB6.64 GiB10.18 GiB0.05 GiB47±12.9%
WizardCoder-Python-13B-V1.0I1-IQ1_S13.0B2.70 GiB6.64 GiB10.18 GiB0.05 GiB47±12.9%
Guanaco-13B-UncensoredI1-IQ1_S13.0B2.70 GiB6.64 GiB10.18 GiB0.05 GiB47±12.9%
QwQ-32BUD-IQ1_S32.8B7.16 GiB2.13 GiB10.18 GiB0.05 GiB47±12.9%
Olmo-3.1-32B-InstructUD-IQ2_XXS32.2B8.30 GiB0.98 GiB10.18 GiB0.05 GiB47±12.9%
Olmo-3.1-32B-ThinkUD-IQ2_XXS32.2B8.30 GiB0.98 GiB10.18 GiB0.05 GiB47±12.9%
Olmo-3-32B-ThinkUD-IQ2_XXS32.2B8.30 GiB0.98 GiB10.18 GiB0.05 GiB47±12.9%
Qwen3-30B-A3BMoEIQ2_S30.5B8.59 GiB0.80 GiB10.18 GiB0.05 GiB129±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation11.62 it/s8.9613.801,506
Benchmarked· n=1,506

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 2080 Ti run?
1679 of 2118 indexed open-weight models fit a GeForce RTX 2080 Ti at 16,384 context with q8_0 KV cache, the largest being Jan-v3-4B-base-instruct at BF16. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 2080 Ti actually have?
Its nameplate is 11 GB, but about 10.23 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 2080 Ti fast for local AI?
Its memory bandwidth is 616 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.