NVIDIA · consumer

GeForce RTX 4080 Laptop

GeForce RTX 4080 Laptop has 12 GB of VRAM at 432 GB/s — about 11.16 GiB usable after driver and compositor overhead. 1616 of 2118 indexed models fit at 64K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
12 GB
GDDR6
Bandwidth
432 GB/s
192-bit bus
Tensor FP16
dense
TDP
150 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1379video 14vision language 135embedding 26image 2audio tts 21audio asr 39

What fits at 64K context

largest quantization that fits, per model · 1616 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-ThinkingQ8_09.4B9.76 GiB0.56 GiB11.16 GiB0.00 GiB30±12.9%
gemma-7bI1-IQ2_XXS8.5B2.41 GiB7.88 GiB11.16 GiB0.00 GiB30±12.9%
Wan2.1-FLF2V-14B-720PQ4_116.4B10.32 GiB0.00 GiB11.16 GiB0.00 GiB30±12.9%
Wan2.1-I2V-14B-480PQ4_116.4B10.32 GiB0.00 GiB11.15 GiB0.01 GiB30±12.9%
Wan2.1-I2V-14B-720PQ4_116.4B10.32 GiB0.00 GiB11.15 GiB0.01 GiB30±12.9%
Qwen3.6-35B-A3B-REAM-160-ru-agentMoEIQ3_M23.6B9.99 GiB0.35 GiB11.15 GiB0.01 GiB116±37%
North-Mini-Code-1.0MoEIQ2_M30.5B9.82 GiB0.55 GiB11.15 GiB0.01 GiB98±37%
Muse-Glimmer-30BUD-IQ2_XXS29.8B10.01 GiB0.26 GiB11.15 GiB0.01 GiB30±12.9%
reka-flash-3.1I1-Q2_K_S20.9B7.95 GiB2.32 GiB11.15 GiB0.01 GiB30±12.9%
Qwen3-16B-A3BMoEQ4_016.0B8.66 GiB1.69 GiB11.14 GiB0.02 GiB54±37%
spoomplesmaxx-v2.1-30BI1-IQ1_S28.9B5.73 GiB4.50 GiB11.14 GiB0.02 GiB30±12.9%
Huihui-granite-4.1-30b-abliteratedI1-IQ1_S28.9B5.73 GiB4.50 GiB11.14 GiB0.02 GiB30±12.9%
granite-4.1-30b-hereticI1-IQ1_S28.9B5.73 GiB4.50 GiB11.14 GiB0.02 GiB30±12.9%
DeepSeek-R1-Distill-Llama-8B-AbliteratedI1-Q3_K_L8.0B8.05 GiB2.25 GiB11.14 GiB0.02 GiB30±12.9%
Nexa-AI-4x4B-InstructMoEI1-Q5_K_S12.1B7.80 GiB2.53 GiB11.14 GiB0.02 GiB27±37%
Salience-1.5-FlashMoEI1-IQ2_S31.1B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
Huihui-Qwen3-VL-30B-A3B-Instruct-abliteratedMoEI1-IQ2_S31.1B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
Qwen3-30B-A3B-Gemini-Pro-High-Reasoning-2507-ABLITERATED-UNCENSOREDMoEI1-IQ2_S30.5B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
MiroThinker-v1.0-30BMoEI1-IQ2_S30.5B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
Qwen3-30B-A3B-YOYO-V5MoEI1-IQ2_S30.5B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
Qwen3-30B-A3B-Thinking-2507-Claude-4.5-Sonnet-High-Reasoning-DistillMoEI1-IQ2_S30.5B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
Huihui-Qwen3-30B-A3B-Thinking-2507-abliteratedMoEI1-IQ2_S30.5B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
Huihui-Qwen3-30B-A3B-Instruct-2507-abliteratedMoEI1-IQ2_S30.5B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
Huihui-Qwen3-Coder-30B-A3B-Instruct-abliteratedMoEI1-IQ2_S30.5B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
Qwen3-Coder-30B-A3B-Instruct-RTPurboMoEI1-IQ2_S30.5B8.65 GiB1.69 GiB11.13 GiB0.03 GiB64±37%
reka-flash-3IQ2_M20.9B7.93 GiB2.32 GiB11.12 GiB0.04 GiB30±12.9%
Gemma-4-31B-Isometry-RPI1-IQ1_S32.7B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
Gemma-4-Dark-Gemistry-31BI1-IQ1_S32.7B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
Prosopon-31BI1-IQ1_S32.7B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
Gemma-4-Novelist-Eclipse-31BI1-IQ1_S32.7B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
Giftige-Blume-31B-v1-StyleSwapI1-IQ1_S32.7B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
G4-MeroMero-31B-StyleSwapI1-IQ1_S32.7B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
Gemma-4-31B-StyleTune-heretic-araI1-IQ1_S32.7B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
Pantheon-Reasoning-31B-1.1I1-IQ1_S32.7B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
Gemma-4-31B-StyleTuneI1-IQ1_S32.7B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
Barcenas-StyleTune-31B-FableI1-IQ1_S32.1B7.10 GiB3.14 GiB11.12 GiB0.04 GiB30±12.9%
Kimi-Linear-48B-A3B-InstructMoEIQ1_S49.1B9.77 GiB0.53 GiB11.11 GiB0.05 GiB30±12.9%
Nemotron-3-Embed-8B-BF16Q8_08.0B7.88 GiB2.39 GiB11.11 GiB0.05 GiB30±12.9%
Marco-Mini-InstructMoEI1-Q3_K_L17.3B8.36 GiB1.97 GiB11.11 GiB0.05 GiB63±37%
Skyfall-31B-v4.2-hereticI1-IQ1_S31.4B6.39 GiB3.80 GiB11.11 GiB0.05 GiB30±12.9%
Skyfall-31B-v4.2I1-IQ1_S31.4B6.39 GiB3.80 GiB11.11 GiB0.05 GiB30±12.9%
Wan2.2-Distill-ModelsQ5_114.3B10.27 GiB0.00 GiB11.10 GiB0.06 GiB30±12.9%
gemma-4-19B-A4B-it-INSTRUCT-Heretic-UncensoredMoEI1-IQ4_XS19.0B9.53 GiB0.79 GiB11.10 GiB0.06 GiB30±12.9%
gemma-4-19B-A4B-it-The-DECKARD-Heretic-Uncensored-ThinkingMoEI1-IQ4_XS19.0B9.53 GiB0.79 GiB11.10 GiB0.06 GiB30±12.9%
gemma-4-19b-a4b-it-REAP-hereticMoEI1-IQ4_XS19.0B9.53 GiB0.79 GiB11.10 GiB0.06 GiB30±12.9%
Gemma-4-19BMoEI1-IQ4_XS19.0B9.53 GiB0.79 GiB11.10 GiB0.06 GiB30±12.9%
Bernini-RQ5_114.3B10.26 GiB0.00 GiB11.10 GiB0.06 GiB30±12.9%
SOLAR-10.7B-Instruct-v1.0-uncensoredQ5_010.7B6.89 GiB3.38 GiB11.10 GiB0.06 GiB30±12.9%
SkyReels-V2-DF-14B-540PQ5_114.3B10.27 GiB0.00 GiB11.10 GiB0.06 GiB30±12.9%
Nous-Hermes-2-SOLAR-10.7BQ5_010.7B6.89 GiB3.38 GiB11.10 GiB0.06 GiB30±12.9%
SOLAR-10.7B-Instruct-v1.0I1-Q5_K_S10.7B6.89 GiB3.38 GiB11.10 GiB0.06 GiB30±12.9%
NVIDIA-Nemotron-Nano-9B-v2Q5_K_S8.9B6.32 GiB3.94 GiB11.10 GiB0.06 GiB30±12.9%
openNemo-9B-abliteratedQ5_K_S8.9B6.32 GiB3.94 GiB11.10 GiB0.06 GiB30±12.9%
gemma-4-26B-A4B-itMoEIQ2_S26.5B9.53 GiB0.79 GiB11.10 GiB0.06 GiB30±12.9%
gemma-3-12b-it-vl-Gemini-3-Pro-Preview-Heretic-Uncensored-ThinkingI1-Q6_K12.2B9.00 GiB1.26 GiB11.10 GiB0.06 GiB30±12.9%
gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-ThinkingI1-Q6_K12.2B9.00 GiB1.26 GiB11.10 GiB0.06 GiB30±12.9%
gemma-3-12b-it-ultra-uncensored-hereticQ6_K12.2B9.00 GiB1.26 GiB11.10 GiB0.06 GiB30±12.9%
gemma-3-12b-it-vl-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-Q6_K12.2B9.00 GiB1.26 GiB11.10 GiB0.06 GiB30±12.9%
Floppa-12B-Gemma3-UncensoredI1-Q6_K12.2B9.00 GiB1.26 GiB11.10 GiB0.06 GiB30±12.9%
gemma-3-12b-it-hereticI1-Q6_K12.2B9.00 GiB1.26 GiB11.10 GiB0.06 GiB30±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation13.40 it/s10.2716.50247
Benchmarked· n=247

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 4080 Laptop run?
1616 of 2118 indexed open-weight models fit a GeForce RTX 4080 Laptop at 65,536 context with q4_0 KV cache, the largest being Qwen3.5-9B-Claude-4.6-Opus-Deckard-V4.2-Uncensored-Heretic-Thinking at Q8_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 4080 Laptop actually have?
Its nameplate is 12 GB, but about 11.16 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 4080 Laptop fast for local AI?
Its memory bandwidth is 432 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.