NVIDIA · consumer

GeForce RTX 3080 Ti Laptop

GeForce RTX 3080 Ti Laptop has 16 GB of VRAM at 512 GB/s — about 14.88 GiB usable after driver and compositor overhead. 1846 of 2118 indexed models fit at 32K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
16 GB
GDDR6
Bandwidth
512 GB/s
256-bit bus
Tensor FP16
dense
TDP
150 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1581audio asr 39vision language 162video 15embedding 26image 2audio tts 21

What fits at 32K context

largest quantization that fits, per model · 1846 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
GLM-4.7-Flash-hereticMoEIQ3_M29.9B13.60 GiB0.46 GiB14.88 GiB0.00 GiB96±37%
gemma-4-E4B-uncensoredF167.9B13.92 GiB0.14 GiB14.88 GiB0.00 GiB26±12.9%
gemma-4-E4B-it-qat-heretic_decensoredF167.9B13.92 GiB0.14 GiB14.88 GiB0.00 GiB26±12.9%
gemma-4-E4B-it-QAT-SOMPOA-heresyF167.9B13.92 GiB0.14 GiB14.88 GiB0.00 GiB26±12.9%
gemma-4-E4B-it-hereticBF168.0B13.92 GiB0.14 GiB14.88 GiB0.00 GiB26±12.9%
Tinman-gemma4-companion-mergedBF167.9B13.92 GiB0.14 GiB14.88 GiB0.00 GiB26±12.9%
v6-Finch-14B-HFQ2_K_L14.1B5.45 GiB8.58 GiB14.87 GiB0.01 GiB26±12.9%
Devstral-Small-2-24B-Instruct-2512IQ4_NL24.0B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Voxtral-Small-24B-2507IQ4_NL24.3B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Morax-24B-v2IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Dolphin3.0-R1-Mistral-24BIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Dolphin3.0-Mistral-24BIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Dans-PersonalityEngine-V1.2.0-24bIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Cydonia_VistralIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Mistral-Small-3.2-24B-Instruct-2506IQ4_NL24.0B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Dans-PersonalityEngine-V1.3.0-24bIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Devstral-Small-2507IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Devstral-Small-2505IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
MS3.2-PaintedFantasy-v3-24BIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Magistral-Small-2509IQ4_NL24.0B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Magistral-Small-2507IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Precog-24B-v1IQ4_NL12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Magidonia-24B-v4.3IQ4_NL12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Magidonia-24B-v4.2.0IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
MS-2501-DPE-QwQify-v0.1-24BIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
sarvam-mIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Magistral-Small-2506IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Cydonia-24B-v4.1IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Cydonia-24B-v4IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Mistral-Small-3.1-24B-Instruct-2503IQ4_NL24.0B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Cydonia-24B-v4.3IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Cydonia-24B-v4.2.0IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
MS3.2-24B-Magnum-DiamondIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Mistral-Small-24B-Instruct-2501-abliteratedIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Dolphin-Mistral-24B-Venice-EditionIQ4_NL24.0B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Mistral-Small-24B-Instruct-2501IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Hearthfire-24BIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Mistral-Small-24B-ArliAI-RPMax-v1.4IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Codex-24B-Small-3.2IQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
Harbinger-24BIQ4_NL23.6B12.54 GiB1.41 GiB14.87 GiB0.01 GiB26±12.9%
spoomplesmaxx-v2.1-30BI1-Q3_K_S28.9B11.71 GiB2.25 GiB14.87 GiB0.01 GiB26±12.9%
Huihui-granite-4.1-30b-abliteratedI1-Q3_K_S28.9B11.71 GiB2.25 GiB14.87 GiB0.01 GiB26±12.9%
granite-4.1-30b-hereticI1-Q3_K_S28.9B11.71 GiB2.25 GiB14.87 GiB0.01 GiB26±12.9%
granite-4.1-30bQ3_K_S28.9B11.71 GiB2.25 GiB14.87 GiB0.01 GiB26±12.9%
EuroLLM-22B-Instruct-2512IQ4_NL22.6B12.09 GiB1.90 GiB14.85 GiB0.03 GiB26±12.9%
Qwen3-Coder-REAP-25B-A3BMoEIQ4_NL24.9B13.22 GiB0.84 GiB14.85 GiB0.03 GiB78±37%
Phi-3-mini-4k-instructKV unresolvedQ6_K3.8B10.67 GiB3.38 GiB14.85 GiB0.03 GiB26±12.9%
Slimaki-Tavern-24B-v1.3Q4_023.6B12.52 GiB1.41 GiB14.84 GiB0.04 GiB26±12.9%
mistral-small-3.1-24b-instruct-2503-hfQ4_023.6B12.52 GiB1.41 GiB14.84 GiB0.04 GiB26±12.9%
Gemma-4-31B-StyleTuneIQ3_S32.7B12.22 GiB1.74 GiB14.84 GiB0.04 GiB26±12.9%
North-Mini-Code-1.0MoEQ3_K_L30.5B13.74 GiB0.32 GiB14.84 GiB0.04 GiB102±37%
gpt-oss-20b-hereticMoEQ4_K_S20.9B13.83 GiB0.22 GiB14.84 GiB0.04 GiB76±37%
Seed-OSS-36B-Instruct-biprojected-norm-preserving-abliteratedI1-IQ2_M36.2B11.68 GiB2.25 GiB14.83 GiB0.05 GiB26±12.9%
Seed-OSS-36B-InstructIQ2_M36.2B11.68 GiB2.25 GiB14.83 GiB0.05 GiB26±12.9%
Hermes-4.3-36B-hereticI1-IQ2_M36.2B11.68 GiB2.25 GiB14.83 GiB0.05 GiB26±12.9%
Hermes-4.3-36BIQ2_M36.2B11.68 GiB2.25 GiB14.83 GiB0.05 GiB26±12.9%
Darwin-35B-A3B-OpusMoEIQ3_XXS36.0B13.85 GiB0.18 GiB14.83 GiB0.05 GiB136±37%
Aurora-Code-1MoEIQ3_XXS34.7B13.85 GiB0.18 GiB14.83 GiB0.05 GiB136±37%
grug-35b-v2MoEIQ3_XXS35.1B13.85 GiB0.18 GiB14.83 GiB0.05 GiB136±37%
grug-35bMoEIQ3_XXS35.1B13.85 GiB0.18 GiB14.83 GiB0.05 GiB136±37%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation9.14 it/s6.4211.63106
Benchmarked· n=106

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 3080 Ti Laptop run?
1846 of 2118 indexed open-weight models fit a GeForce RTX 3080 Ti Laptop at 32,768 context with q4_0 KV cache, the largest being GLM-4.7-Flash-heretic at IQ3_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3080 Ti Laptop actually have?
Its nameplate is 16 GB, but about 14.88 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3080 Ti Laptop fast for local AI?
Its memory bandwidth is 512 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.