NVIDIA · consumer

GeForce GTX 1080 Ti

GeForce GTX 1080 Ti has 11 GB of VRAM at 484 GB/s — about 10.23 GiB usable after driver and compositor overhead. 1670 of 2118 indexed models fit at 32K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
11 GB
GDDR5X
Bandwidth
484 GB/s
352-bit bus
Tensor FP16
dense
TDP
250 W
$699 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1437vision language 131video 14audio tts 21image 2audio asr 39embedding 26

What fits at 32K context

largest quantization that fits, per model · 1670 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3-30B-A3BMoEIQ2_S30.5B8.59 GiB0.84 GiB10.23 GiB0.00 GiB101±37%
Pantheon-Proto-RP-1.8-30B-A3BMoEIQ2_S30.5B8.59 GiB0.84 GiB10.23 GiB0.00 GiB101±37%
NousCoder-14BQ4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
spoomplesmaxx-mini-14BI1-Q4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
vanilla-cn-roleplay-0.2I1-Q4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
Claria-14bI1-Q4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
NTX-2.1-ProI1-Q4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
Qwen3-14B-UncensoredI1-Q4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
Qwen3-14BQ4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
FrogMini-14B-2510I1-Q4_07.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
Qwen3-14B-abliteratedQ4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
Josiefied-Qwen3-14B-abliterated-v3Q4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
Hermes-4-14BQ4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
Slava-Qwen3-14B-SerbianI1-Q4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
Huihui-Qwen3-14B-abliterated-v2I1-Q4_014.8B7.96 GiB1.41 GiB10.22 GiB0.01 GiB37±12.9%
starcoder2-15bKV unresolvedQ4_K_S16.0B8.62 GiB0.70 GiB10.22 GiB0.01 GiB37±12.9%
EuroLLM-22B-Instruct-2512IQ2_M22.6B7.45 GiB1.90 GiB10.21 GiB0.02 GiB37±12.9%
gemma-4-26B-A4B-itMoEIQ2_XXS26.5B8.99 GiB0.43 GiB10.21 GiB0.02 GiB37±12.9%
glm-4-9b-chat-1mQ2_K9.5B3.74 GiB5.63 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-16B-A3BMoEIQ4_NL16.0B8.58 GiB0.84 GiB10.21 GiB0.02 GiB81±37%
next-8bQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
next-ocrQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-VL-8B-GLM-4.7-Flash-Heretic-Uncensored-ThinkingQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Midas-FableAgent-8BQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-VL-8B-Heretic-1.3.0Q8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-VL-8B-ThinkingQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-VL-8B-Instruct-Unredacted-MAXQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Poe-8B-GLM5-Opus4.6-Sonnet4.5-Kimi-Grok-Gemini-3-pro-preview-HERETICQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen-3-VL-8B-Instruct-hereticQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
nsfwcaption-qwen3-vl-8b-v3-safetensorsQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Huihui-Qwen3-VL-8B-Instruct-abliteratedQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-VL-Reranker-8BQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Salience-1-9BQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-VL-8B-InstructQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-VL-8B-Instruct-Uncensored-V2Q8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
GRaPE-2-FlashQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Jan-v2-VL-highQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Jan-v2-VL-medQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
DeepSeek-R1-0528-Qwen3-8BQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Parable-Qwen3-8B-Claude-Fable-5Q8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
ReasonCritic-7BQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Finch-8B-KTOQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Finch-8BQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
mythos-9b-unhinged-hereticQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
nsfwvision-qwen3-vl-8b-v3-safetensorsQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
MathSmith-hc-Qwen3-8BQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-VL-8B-Thinking-Unredacted-MAXQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
MiroThinker-v1.0-8BQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
mythos-9b-unhingedQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Maestro1-9BQ8_08.8B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
qwen3-8b-claude-agentic-fable5Q8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Ektome-Qwen3-8B-PristinelyUncensoredQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
mythos-9b-mergedQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-8BQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
qwen3-8b-apostateQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Josiefied-Qwen3-8B-abliterated-v1Q8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
tmax-8bQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-8B-abliteratedQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Marco-DeepResearch-8BQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
Qwen3-8B-UncensoredQ8_08.2B8.11 GiB1.27 GiB10.21 GiB0.02 GiB37±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.19 it/s2.123.64422
Benchmarked· n=422

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce GTX 1080 Ti run?
1670 of 2118 indexed open-weight models fit a GeForce GTX 1080 Ti at 32,768 context with q4_0 KV cache, the largest being Qwen3-30B-A3B at IQ2_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce GTX 1080 Ti actually have?
Its nameplate is 11 GB, but about 10.23 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce GTX 1080 Ti fast for local AI?
Its memory bandwidth is 484 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.