GeForce RTX 3080 Ti Laptop
GeForce RTX 3080 Ti Laptop has 16 GB of VRAM at 512 GB/s — about 14.88 GiB usable after driver and compositor overhead. 1846 of 2118 indexed models fit at 32K context with q4_0 KV.
What fits at 32K context
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| GLM-4.7-Flash-hereticMoE | IQ3_M | 29.9B | 13.60 GiB | 0.46 GiB | 14.88 GiB | 0.00 GiB | 96±37% |
| gemma-4-E4B-uncensored | F16 | 7.9B | 13.92 GiB | 0.14 GiB | 14.88 GiB | 0.00 GiB | 26±12.9% |
| gemma-4-E4B-it-qat-heretic_decensored | F16 | 7.9B | 13.92 GiB | 0.14 GiB | 14.88 GiB | 0.00 GiB | 26±12.9% |
| gemma-4-E4B-it-QAT-SOMPOA-heresy | F16 | 7.9B | 13.92 GiB | 0.14 GiB | 14.88 GiB | 0.00 GiB | 26±12.9% |
| gemma-4-E4B-it-heretic | BF16 | 8.0B | 13.92 GiB | 0.14 GiB | 14.88 GiB | 0.00 GiB | 26±12.9% |
| Tinman-gemma4-companion-merged | BF16 | 7.9B | 13.92 GiB | 0.14 GiB | 14.88 GiB | 0.00 GiB | 26±12.9% |
| v6-Finch-14B-HF | Q2_K_L | 14.1B | 5.45 GiB | 8.58 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Devstral-Small-2-24B-Instruct-2512 | IQ4_NL | 24.0B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Voxtral-Small-24B-2507 | IQ4_NL | 24.3B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Morax-24B-v2 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Dolphin3.0-R1-Mistral-24B | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Dolphin3.0-Mistral-24B | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Dans-PersonalityEngine-V1.2.0-24b | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Cydonia_Vistral | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Mistral-Small-3.2-24B-Instruct-2506 | IQ4_NL | 24.0B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Dans-PersonalityEngine-V1.3.0-24b | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Devstral-Small-2507 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Devstral-Small-2505 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| MS3.2-PaintedFantasy-v3-24B | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Magistral-Small-2509 | IQ4_NL | 24.0B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Magistral-Small-2507 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Precog-24B-v1 | IQ4_NL | — | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Magidonia-24B-v4.3 | IQ4_NL | — | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Magidonia-24B-v4.2.0 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| MS-2501-DPE-QwQify-v0.1-24B | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| sarvam-m | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Magistral-Small-2506 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Cydonia-24B-v4.1 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Cydonia-24B-v4 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Mistral-Small-3.1-24B-Instruct-2503 | IQ4_NL | 24.0B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Cydonia-24B-v4.3 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Cydonia-24B-v4.2.0 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| MS3.2-24B-Magnum-Diamond | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Mistral-Small-24B-Instruct-2501-abliterated | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Dolphin-Mistral-24B-Venice-Edition | IQ4_NL | 24.0B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Mistral-Small-24B-Instruct-2501 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Hearthfire-24B | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Mistral-Small-24B-ArliAI-RPMax-v1.4 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Codex-24B-Small-3.2 | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Harbinger-24B | IQ4_NL | 23.6B | 12.54 GiB | 1.41 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| spoomplesmaxx-v2.1-30B | I1-Q3_K_S | 28.9B | 11.71 GiB | 2.25 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| Huihui-granite-4.1-30b-abliterated | I1-Q3_K_S | 28.9B | 11.71 GiB | 2.25 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| granite-4.1-30b-heretic | I1-Q3_K_S | 28.9B | 11.71 GiB | 2.25 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| granite-4.1-30b | Q3_K_S | 28.9B | 11.71 GiB | 2.25 GiB | 14.87 GiB | 0.01 GiB | 26±12.9% |
| EuroLLM-22B-Instruct-2512 | IQ4_NL | 22.6B | 12.09 GiB | 1.90 GiB | 14.85 GiB | 0.03 GiB | 26±12.9% |
| Qwen3-Coder-REAP-25B-A3BMoE | IQ4_NL | 24.9B | 13.22 GiB | 0.84 GiB | 14.85 GiB | 0.03 GiB | 78±37% |
| Phi-3-mini-4k-instructKV unresolved | Q6_K | 3.8B | 10.67 GiB | 3.38 GiB | 14.85 GiB | 0.03 GiB | 26±12.9% |
| Slimaki-Tavern-24B-v1.3 | Q4_0 | 23.6B | 12.52 GiB | 1.41 GiB | 14.84 GiB | 0.04 GiB | 26±12.9% |
| mistral-small-3.1-24b-instruct-2503-hf | Q4_0 | 23.6B | 12.52 GiB | 1.41 GiB | 14.84 GiB | 0.04 GiB | 26±12.9% |
| Gemma-4-31B-StyleTune | IQ3_S | 32.7B | 12.22 GiB | 1.74 GiB | 14.84 GiB | 0.04 GiB | 26±12.9% |
| North-Mini-Code-1.0MoE | Q3_K_L | 30.5B | 13.74 GiB | 0.32 GiB | 14.84 GiB | 0.04 GiB | 102±37% |
| gpt-oss-20b-hereticMoE | Q4_K_S | 20.9B | 13.83 GiB | 0.22 GiB | 14.84 GiB | 0.04 GiB | 76±37% |
| Seed-OSS-36B-Instruct-biprojected-norm-preserving-abliterated | I1-IQ2_M | 36.2B | 11.68 GiB | 2.25 GiB | 14.83 GiB | 0.05 GiB | 26±12.9% |
| Seed-OSS-36B-Instruct | IQ2_M | 36.2B | 11.68 GiB | 2.25 GiB | 14.83 GiB | 0.05 GiB | 26±12.9% |
| Hermes-4.3-36B-heretic | I1-IQ2_M | 36.2B | 11.68 GiB | 2.25 GiB | 14.83 GiB | 0.05 GiB | 26±12.9% |
| Hermes-4.3-36B | IQ2_M | 36.2B | 11.68 GiB | 2.25 GiB | 14.83 GiB | 0.05 GiB | 26±12.9% |
| Darwin-35B-A3B-OpusMoE | IQ3_XXS | 36.0B | 13.85 GiB | 0.18 GiB | 14.83 GiB | 0.05 GiB | 136±37% |
| Aurora-Code-1MoE | IQ3_XXS | 34.7B | 13.85 GiB | 0.18 GiB | 14.83 GiB | 0.05 GiB | 136±37% |
| grug-35b-v2MoE | IQ3_XXS | 35.1B | 13.85 GiB | 0.18 GiB | 14.83 GiB | 0.05 GiB | 136±37% |
| grug-35bMoE | IQ3_XXS | 35.1B | 13.85 GiB | 0.18 GiB | 14.83 GiB | 0.05 GiB | 136±37% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Measured on this card
| Workload◍ | Median | Middle 50% | Runs |
|---|---|---|---|
| Image generation | 9.14 it/s | 6.42–11.63 | 106 |
Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.
Questions people ask
- What AI models can a GeForce RTX 3080 Ti Laptop run?
- 1846 of 2118 indexed open-weight models fit a GeForce RTX 3080 Ti Laptop at 32,768 context with q4_0 KV cache, the largest being GLM-4.7-Flash-heretic at IQ3_M. That covers text, vision-language, image, video and speech models.
- How much usable memory does a GeForce RTX 3080 Ti Laptop actually have?
- Its nameplate is 16 GB, but about 14.88 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a GeForce RTX 3080 Ti Laptop fast for local AI?
- Its memory bandwidth is 512 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.