NVIDIA · consumer

GeForce GTX 1080 Ti

GeForce GTX 1080 Ti has 11 GB of VRAM at 484 GB/s — about 10.23 GiB usable after driver and compositor overhead. 1502 of 2118 indexed models fit at 32K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
11 GB
GDDR5X
Bandwidth
484 GB/s
352-bit bus
Tensor FP16
dense
TDP
250 W
$699 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1283vision language 119video 14audio tts 21image 1audio asr 38embedding 26

What fits at 32K context

largest quantization that fits, per model · 1502 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Ministral-3-14B-Instruct-2512-BF16-abliteratedI1-Q3_K_L13.9B6.72 GiB2.66 GiB10.23 GiB0.00 GiB37±12.9%
Ministral-3-14B-abliteratedQ3_K_L13.9B6.72 GiB2.66 GiB10.23 GiB0.00 GiB37±12.9%
Ministral-3-14B-Instruct-2512-BF16Q3_K_L13.9B6.72 GiB2.66 GiB10.23 GiB0.00 GiB37±12.9%
Ministral-3-14B-Reasoning-2512-UncensoredI1-Q3_K_L13.9B6.72 GiB2.66 GiB10.23 GiB0.00 GiB37±12.9%
dolphincoder-starcoder2-15bKV unresolvedI1-IQ4_XS16.0B8.01 GiB1.33 GiB10.23 GiB0.00 GiB37±12.9%
starcoder2-15bKV unresolvedIQ4_XS16.0B8.01 GiB1.33 GiB10.23 GiB0.00 GiB37±12.9%
Smilodon-9B-v1I1-Q5_K_M10.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
bella-bartender-v2I1-Q5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-Q5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-Q5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
Gemma-2-9B-It-SPPO-Iter3I1-Q5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
Gemma-SEA-LION-v3-9B-ITI1-Q5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-Q5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
G2-Darkest-Writer-9B-v1I1-Q5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
Tiger-Gemma-9B-v3I1-Q5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
gemma-2-9b-it-abliteratedQ5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
gemma-2-9b-itQ5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
Tiger-Gemma-9B-v1Q5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
magnum-v4-9bQ5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
gemma-2-9bQ5_K_M9.2B6.19 GiB3.18 GiB10.21 GiB0.02 GiB37±12.9%
granite-4.1-8bQ6_K8.8B6.72 GiB2.66 GiB10.21 GiB0.02 GiB37±12.9%
Phi-3-medium-128k-instructIQ3_M14.0B6.03 GiB3.32 GiB10.21 GiB0.02 GiB37±12.9%
Phi-3-medium-4k-instructI1-IQ3_M14.0B6.03 GiB3.32 GiB10.21 GiB0.02 GiB37±12.9%
Rocinante-XL-16B-v1I1-Q2_K16.1B5.77 GiB3.59 GiB10.20 GiB0.03 GiB37±12.9%
Llama3.2-24B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEI1-IQ3_S18.0B7.53 GiB1.86 GiB10.20 GiB0.03 GiB61±37%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopQ6_K12.1B8.93 GiB0.40 GiB10.19 GiB0.04 GiB37±12.9%
DeepSeek-Coder-V2-Lite-BaseMoEI1-Q4_K_S15.7B8.88 GiB0.50 GiB10.19 GiB0.04 GiB103±37%
DeepSeek-Coder-V2-Lite-InstructMoEQ4_K_S15.7B8.88 GiB0.50 GiB10.19 GiB0.04 GiB103±37%
DeepSeek-V2-Lite-ChatMoEQ4_K_S15.7B8.88 GiB0.50 GiB10.19 GiB0.04 GiB103±37%
gemma-4-A4B-98e-v6-coder-itMoEIQ3_XS20.5B8.58 GiB0.82 GiB10.19 GiB0.04 GiB37±12.9%
DeepSeek-V2-Lite-Chat-Uncensored-Unbiased-ReasonerMoEQ4_K_S15.7B8.88 GiB0.50 GiB10.19 GiB0.04 GiB103±37%
DeepSeek-V2-Lite-Chat-UncensoredMoEQ4_K_S15.7B8.87 GiB0.50 GiB10.19 GiB0.04 GiB103±37%
Fimbulvetr-11B-v2I1-Q4_K_M10.7B6.16 GiB3.19 GiB10.19 GiB0.04 GiB37±12.9%
Tiger-Gemma-12B-v3Q4_K_L12.8B8.02 GiB1.31 GiB10.18 GiB0.05 GiB37±12.9%
Falcon3-10B-InstructQ5_K_S10.3B6.65 GiB2.66 GiB10.18 GiB0.05 GiB37±12.9%
Snowpiercer-15B-v4Q2_K_L15.0B6.01 GiB3.32 GiB10.17 GiB0.06 GiB37±12.9%
NVIDIA-Nemotron-Nano-12B-v2Q3_K_S12.3B5.19 GiB4.12 GiB10.17 GiB0.06 GiB37±12.9%
Ling-liteMoEQ3_K_L16.8B8.45 GiB0.93 GiB10.17 GiB0.06 GiB87±37%
Mistral-NeMo-Minitron-8B-InstructQ6_K_L8.4B6.68 GiB2.66 GiB10.16 GiB0.07 GiB37±12.9%
Cydonia-v1.3-Magnum-v4-22BI1-IQ2_XXS22.2B5.58 GiB3.72 GiB10.16 GiB0.07 GiB37±12.9%
Mistral-Small-22B-ArliAI-RPMax-v1.1I1-IQ2_XXS22.2B5.58 GiB3.72 GiB10.16 GiB0.07 GiB37±12.9%
magnum-v4-22bI1-IQ2_XXS22.2B5.58 GiB3.72 GiB10.16 GiB0.07 GiB37±12.9%
Mistral-Small-Instruct-2409IQ2_XXS22.2B5.58 GiB3.72 GiB10.16 GiB0.07 GiB37±12.9%
Codestral-22B-v0.1IQ2_XXS22.2B5.58 GiB3.72 GiB10.16 GiB0.07 GiB37±12.9%
dolphin-2.9.1-mixtral-1x22bMoEI1-IQ2_XXS22.2B5.58 GiB3.72 GiB10.16 GiB0.07 GiB21±37%
glm-4v-9bQ8_013.9B9.31 GiB0.00 GiB10.16 GiB0.07 GiB37±12.9%
Snowpiercer-15B-v4-hereticI1-IQ3_XS15.0B5.98 GiB3.32 GiB10.15 GiB0.08 GiB37±12.9%
gemma-4-12BQ5_012.0B7.99 GiB1.31 GiB10.15 GiB0.08 GiB37±12.9%
North-Mini-Code-1.0MoEIQ2_XS30.5B8.77 GiB0.60 GiB10.15 GiB0.08 GiB113±37%
NuExtract-1.5Q6_K_L3.8B2.96 GiB6.38 GiB10.15 GiB0.08 GiB37±12.9%
Phi-3.5-mini-instructQ6_K_L3.8B2.96 GiB6.38 GiB10.15 GiB0.08 GiB37±12.9%
Phi-3-mini-128k-instructQ6_K_L3.8B2.96 GiB6.38 GiB10.15 GiB0.08 GiB37±12.9%
Phi-3.5-mini-instruct_UncensoredQ6_K_L3.8B2.96 GiB6.38 GiB10.15 GiB0.08 GiB37±12.9%
Phi-3-mini-4k-instructQ6_K_L3.8B2.96 GiB6.38 GiB10.15 GiB0.08 GiB37±12.9%
Skywork-R1V3-38BIQ2_XS38.4B9.27 GiB0.00 GiB10.14 GiB0.09 GiB37±12.9%
L3.2-Rogue-Creative-Instruct-Uncensored-Abliterated-7BQ5_K_S7.5B4.88 GiB4.45 GiB10.14 GiB0.09 GiB37±12.9%
gemma-4-12B-coder-fable5-composer2.5-v1-abliteratedQ4_K_M12.0B7.98 GiB1.31 GiB10.14 GiB0.09 GiB37±12.9%
gemma-4-12B-coder-fable5-composer2.5-v1-sft-v5-abliteratedQ4_K_M12.0B7.98 GiB1.31 GiB10.14 GiB0.09 GiB37±12.9%
dolphin-2.9.3-mistral-7B-32kQ8_07.2B7.17 GiB2.13 GiB10.14 GiB0.09 GiB37±12.9%
Mistral-7B-v0.3Q8_07.2B7.17 GiB2.13 GiB10.14 GiB0.09 GiB37±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation3.19 it/s2.123.64422
Benchmarked· n=422

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce GTX 1080 Ti run?
1502 of 2118 indexed open-weight models fit a GeForce GTX 1080 Ti at 32,768 context with q8_0 KV cache, the largest being Ministral-3-14B-Instruct-2512-BF16-abliterated at I1-Q3_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce GTX 1080 Ti actually have?
Its nameplate is 11 GB, but about 10.23 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce GTX 1080 Ti fast for local AI?
Its memory bandwidth is 484 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.