NVIDIA · consumer

GeForce RTX 2080 Ti

GeForce RTX 2080 Ti has 11 GB of VRAM at 616 GB/s — about 10.23 GiB usable after driver and compositor overhead. 1714 of 2118 indexed models fit at 8K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
11 GB
GDDR6
Bandwidth
616 GB/s
352-bit bus
Tensor FP16
108 TF
dense
TDP
250 W
$999 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1474audio asr 39vision language 138video 14audio tts 21image 2embedding 26

What fits at 8K context

largest quantization that fits, per model · 1714 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Phi-3-mini-4k-instructKV unresolvedQ4_K_M3.8B7.83 GiB1.59 GiB10.23 GiB0.00 GiB46±12.9%
Voxtral-Small-24B-2507IQ3_XXS24.3B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Devstral-Small-2-24B-Instruct-2512IQ3_XXS24.0B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
ERNIE-4.5-21B-A3B-ThinkingQ3_K_S21.8B9.17 GiB0.23 GiB10.23 GiB0.00 GiB46±12.9%
ERNIE-4.5-21B-A3B-PTQ3_K_S21.9B9.17 GiB0.23 GiB10.23 GiB0.00 GiB46±12.9%
Transformed-Journey-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Magistry-24B-v1.1I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Mergedonia-AETHER-24B-v1aI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Mergedonia-AETHER-24B-v1bI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Slimaki-Tavern-24B-v1.3I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Maginum-Cydoms-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Maginum-Cydoms-24B-absolute-heresyI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Morax-24B-v2IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Mistral-Small-3.2-24B-Instruct-2506-ultra-uncensored-hereticI1-IQ3_XXS24.0B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Huihui-Mistral-Small-3.2-24B-Instruct-2506-abliterated-llamacppfixedI1-IQ3_XXS24.0B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Dans-PersonalityEngine-V1.2.0-24bI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Mistral-Small-3_2-24B-Instruct-2506-antislop.v2I1-IQ3_XXS24.0B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Dans-PersonalityEngine-V1.3.0-24bI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia_VistralIQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Dolphin3.0-Mistral-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Goetia-24B-v1.1I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Devstral-Small-2505IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Mistral-Small-3.2-24B-Instruct-2506IQ3_XXS24.0B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
MS3.2-PaintedFantasy-v3-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
RP-Spectrum-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
MS3.2-PaintedFantasy-v4.1-24B-ultra-uncensored-heretic-v2I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Magidonia-24B-v4.3-heretic-v1.2I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Magidonia-24B-v4.3-absolute-heresyI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
MagiSeek-Pro-V1I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cogidonia-v2-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Magidonia-24B-v4.3I1-IQ3_XXS8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Precog-24B-v1I1-IQ3_XXS8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
experiment024bI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Magidonia-24B-v4.2.0IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Berthier-Mistral-Military-24BI1-IQ3_XXS24.0B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
MS-2501-DPE-QwQify-v0.1-24BIQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Mistral-Small-3.2-24B-Instruct-2506-llamacppfixedI1-IQ3_XXS24.0B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia-24B-v4.3-absolute-heresyI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia-24B-v4.3-heretic-v2I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia-24B-v4.3-hereticI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia-24B-v4.3-heretic-v4I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia-24B-v4.2.0I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Journeys-End-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
sarvam-mIQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Dolphin-Mistral-GLM-4.7-Flash-24B-Venice-Edition-Thinking-UncensoredI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
WeirdCompound-v1.7-24bI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Magistral-Small-2506IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia-24B-v4.3I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia-24B-v4.1IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia-24B-v4IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Mistral-Small-3.1-24B-Instruct-2503IQ3_XXS24.0B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Mistral-Small-24B-Instruct-JbliteratedI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Mistral-Small-24B-Instruct-2501-abliteratedI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
MS3.2-24B-Magnum-DiamondIQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Dolphin-Mistral-24B-Venice-EditionI1-IQ3_XXS24.0B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
WeirdDolphinPersonalityMechanism-Mistral-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
grok-oss-Apollyon-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
grok-oss-Apollyon-24B-hereticI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Cydonia-v4.1-MS3.2-Magnum-Diamond-24BI1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
Codex-24B-Small-3.2I1-IQ3_XXS23.6B8.64 GiB0.66 GiB10.23 GiB0.00 GiB47±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation11.62 it/s8.9613.801,506
Benchmarked· n=1,506

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 2080 Ti run?
1714 of 2118 indexed open-weight models fit a GeForce RTX 2080 Ti at 8,192 context with q8_0 KV cache, the largest being Phi-3-mini-4k-instruct at Q4_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 2080 Ti actually have?
Its nameplate is 11 GB, but about 10.23 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 2080 Ti fast for local AI?
Its memory bandwidth is 616 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.