NVIDIA · consumer

GeForce GTX 1080 Ti

GeForce GTX 1080 Ti has 11 GB of VRAM at 484 GB/s — about 10.23 GiB usable after driver and compositor overhead. 1501 of 2118 indexed models fit at 64K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
11 GB
GDDR5X
Bandwidth
484 GB/s
352-bit bus
Tensor FP16
dense
TDP
250 W
$699 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1282vision language 119video 14audio tts 21image 1audio asr 38embedding 26

What fits at 64K context

largest quantization that fits, per model · 1501 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Nous-Hermes-2-SOLAR-10.7BQ4_K_M10.7B6.02 GiB3.38 GiB10.23 GiB0.00 GiB37±12.9%
SOLAR-10.7B-Instruct-v1.0I1-Q4_K_M10.7B6.02 GiB3.38 GiB10.23 GiB0.00 GiB37±12.9%
Qwen3-Coder-REAP-25B-A3BMoEIQ2_M24.9B7.75 GiB1.69 GiB10.23 GiB0.00 GiB71±37%
Ling-liteMoEQ3_K_L16.8B8.45 GiB0.98 GiB10.23 GiB0.00 GiB85±37%
DeepSeek-Coder-V2-Lite-BaseMoEI1-Q4_K_S15.7B8.88 GiB0.53 GiB10.22 GiB0.01 GiB101±37%
DeepSeek-Coder-V2-Lite-InstructMoEQ4_K_S15.7B8.88 GiB0.53 GiB10.22 GiB0.01 GiB101±37%
DeepSeek-V2-Lite-ChatMoEQ4_K_S15.7B8.88 GiB0.53 GiB10.22 GiB0.01 GiB101±37%
NVIDIA-Nemotron-Nano-12B-v2Q2_K_L12.3B4.99 GiB4.36 GiB10.22 GiB0.01 GiB37±12.9%
DeepSeek-V2-Lite-Chat-Uncensored-Unbiased-ReasonerMoEQ4_K_S15.7B8.88 GiB0.53 GiB10.22 GiB0.01 GiB102±37%
Nexa-AI-4x4B-InstructMoEI1-Q4_K_M12.1B6.88 GiB2.53 GiB10.22 GiB0.01 GiB32±37%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopQ6_K12.1B8.93 GiB0.42 GiB10.22 GiB0.01 GiB37±12.9%
DeepSeek-V2-Lite-Chat-UncensoredMoEQ4_K_S15.7B8.87 GiB0.53 GiB10.22 GiB0.01 GiB102±37%
Falcon3-7B-InstructQ8_07.5B7.38 GiB1.97 GiB10.22 GiB0.01 GiB37±12.9%
Nemotron-Mini-4B-InstructQ3_K_M4.2B7.15 GiB2.25 GiB10.22 GiB0.01 GiB37±12.9%
Phi-4-mini-instruct-abliteratedF163.8B7.15 GiB2.25 GiB10.21 GiB0.02 GiB37±12.9%
Phi-4-mini-reasoningBF163.8B7.15 GiB2.25 GiB10.21 GiB0.02 GiB37±12.9%
Phi-4-mini-instructBF163.8B7.15 GiB2.25 GiB10.21 GiB0.02 GiB37±12.9%
NVIDIA-Nemotron-Nano-9B-v2Q4_18.9B5.43 GiB3.94 GiB10.21 GiB0.02 GiB37±12.9%
Ministral-3-8B-Instruct-2512-BF16Q6_K_M8.9B6.98 GiB2.39 GiB10.21 GiB0.02 GiB37±12.9%
Fallen-Gemma3-27B-v1IQ2_S27.4B8.18 GiB1.18 GiB10.20 GiB0.03 GiB37±12.9%
gemma-3-12b-it-abliteratedQ5_K_L12.2B8.09 GiB1.26 GiB10.19 GiB0.04 GiB37±12.9%
Phi-4-reasoning-plusIQ3_XS14.7B5.82 GiB3.52 GiB10.19 GiB0.04 GiB37±12.9%
Phi-4-reasoningIQ3_XS14.7B5.82 GiB3.52 GiB10.19 GiB0.04 GiB37±12.9%
phi-4IQ3_XS14.7B5.82 GiB3.52 GiB10.19 GiB0.04 GiB37±12.9%
Smilodon-9B-v1I1-Q5_K_M10.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
bella-bartender-v2I1-Q5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-Q5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-Q5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
Gemma-2-9B-It-SPPO-Iter3I1-Q5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
Gemma-SEA-LION-v3-9B-ITI1-Q5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-Q5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
G2-Darkest-Writer-9B-v1I1-Q5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
Tiger-Gemma-9B-v3I1-Q5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
gemma-2-9b-it-abliteratedQ5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
gemma-2-9b-itQ5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
Tiger-Gemma-9B-v1Q5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
magnum-v4-9bQ5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
gemma-2-9bQ5_K_M9.2B6.19 GiB3.16 GiB10.19 GiB0.04 GiB37±12.9%
Llama3.2-30B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEI1-IQ1_S30.0B6.07 GiB3.30 GiB10.19 GiB0.04 GiB45±37%
ERNIE-4.5-21B-A3B-ThinkingIQ3_XXS21.8B8.38 GiB0.98 GiB10.18 GiB0.05 GiB37±12.9%
ERNIE-4.5-21B-A3B-PTIQ3_XXS21.9B8.38 GiB0.98 GiB10.18 GiB0.05 GiB37±12.9%
NuExtract-1.5Q5_K_M3.8B2.62 GiB6.75 GiB10.18 GiB0.05 GiB37±12.9%
Phi-3.5-mini-instructQ5_K_M3.8B2.62 GiB6.75 GiB10.18 GiB0.05 GiB37±12.9%
Phi-3.5-mini-instruct_UncensoredQ5_K_M3.8B2.62 GiB6.75 GiB10.18 GiB0.05 GiB37±12.9%
Phi-3-mini-128k-instructQ5_K_M3.8B2.62 GiB6.75 GiB10.18 GiB0.05 GiB37±12.9%
Phi-3-mini-4k-instructQ5_K_M3.8B2.62 GiB6.75 GiB10.18 GiB0.05 GiB37±12.9%
octo-netQ5_K3.8B2.62 GiB6.75 GiB10.18 GiB0.05 GiB37±12.9%
ERNIE-21B-A3B-Thinking-Gemini-3-Pro-High-Reasoning-V2I1-IQ3_XS21.8B8.37 GiB0.98 GiB10.17 GiB0.06 GiB37±12.9%
ERNIE-21B-A3B-Claude-4.5-High-OPUS-ThinkingI1-IQ3_XS21.8B8.37 GiB0.98 GiB10.17 GiB0.06 GiB37±12.9%
EVA-abliterated-TIES-Qwen2.5-14BI1-IQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
Neuron-V1-14B-InstructI1-IQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
Ektome-Qwen2.5-Coder-14B-Instruct-PristinelyUncensoredI1-IQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
Qwen2.5-14B-Instruct-1M-abliteratedI1-IQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
DeepCoder-14B-PreviewIQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
Deepseeker-Kunou-Qwen2.5-14bI1-IQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
SuperNova-MediusIQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
14B-Qwen2.5-Kunou-v1I1-IQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
Sugoi-14B-Ultra-HFI1-IQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
Qwen2.5-Coder-14B-Instruct-abliteratedIQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
OpenCodeReasoning-Nemotron-14BIQ3_XS14.8B5.94 GiB3.38 GiB10.17 GiB0.06 GiB37±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.19 it/s2.123.64422
Benchmarked· n=422

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce GTX 1080 Ti run?
1501 of 2118 indexed open-weight models fit a GeForce GTX 1080 Ti at 65,536 context with q4_0 KV cache, the largest being Nous-Hermes-2-SOLAR-10.7B at Q4_K_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce GTX 1080 Ti actually have?
Its nameplate is 11 GB, but about 10.23 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce GTX 1080 Ti fast for local AI?
Its memory bandwidth is 484 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.