GeForce RTX 2070 SUPER
GeForce RTX 2070 SUPER has 8 GB of VRAM at 448 GB/s — about 7.44 GiB usable after driver and compositor overhead. 919 of 2118 indexed models fit at 128K context with q4_0 KV.
What fits at 128K context
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| OLMoE-1B-7B-0924-InstructMoE | I1-IQ2_M | 6.9B | 2.17 GiB | 4.50 GiB | 7.44 GiB | 0.00 GiB | 37±37% |
| CycleGRPO-4B | I1-IQ2_M | 4.8B | 1.57 GiB | 5.06 GiB | 7.44 GiB | 0.00 GiB | 48±12.9% |
| Jan-v3-4B-base-instruct | IQ2_M | 4.4B | 1.56 GiB | 5.06 GiB | 7.44 GiB | 0.00 GiB | 48±12.9% |
| Jan-code-4b | IQ2_M | 4.4B | 1.56 GiB | 5.06 GiB | 7.44 GiB | 0.00 GiB | 48±12.9% |
| OmniAtlas-Qwen3-30B-A3B | I1-IQ1_M | 31.7B | 6.59 GiB | 0.00 GiB | 7.44 GiB | 0.00 GiB | 48±12.9% |
| Qwen3-Omni-30B-A3B-Captioner | I1-IQ1_M | 31.7B | 6.59 GiB | 0.00 GiB | 7.44 GiB | 0.00 GiB | 48±12.9% |
| DeepSeek-Coder-V2-Lite-BaseMoE | I1-IQ2_XS | 15.7B | 5.56 GiB | 1.07 GiB | 7.44 GiB | 0.00 GiB | 92±37% |
| DeepSeek-Coder-V2-Lite-InstructMoE | IQ2_XS | 15.7B | 5.56 GiB | 1.07 GiB | 7.44 GiB | 0.00 GiB | 92±37% |
| DeepSeek-V2-Lite-ChatMoE | IQ2_XS | 15.7B | 5.56 GiB | 1.07 GiB | 7.44 GiB | 0.00 GiB | 92±37% |
| EVA-Yi-1.5-9B-32K-V1 | I1-IQ3_XXS | 8.8B | 3.24 GiB | 3.38 GiB | 7.44 GiB | 0.00 GiB | 48±12.9% |
| Nemotron-3-Embed-8B-BF16 | IQ1_S | 8.0B | 1.81 GiB | 4.78 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-VL-4B-Thinking | UD-IQ3_XXS | 4.4B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-VL-4B-Instruct | UD-IQ3_XXS | 4.4B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-4B | UD-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-4B-Thinking-2507 | UD-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Jan-nano-128k | UD-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-4B-Instruct-2507 | UD-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Jan-nano | UD-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Gemma-4-12B-StyleTune | I1-IQ2_S | 13.0B | 4.20 GiB | 2.38 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| gemma-4-12b-heretic-styletune-head | I1-IQ2_S | 12.0B | 4.20 GiB | 2.38 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| syrian-gemma-12b | I1-IQ2_S | 13.0B | 4.20 GiB | 2.38 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| orpheus-3b-0.1-pretrained | Q5_1 | 3.8B | 2.68 GiB | 3.94 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-VL-4B-Instruct-Unredacted-MAX | I1-IQ3_XXS | 4.4B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-VL-4B-Thinking-Unredacted-MAX | I1-IQ3_XXS | 4.4B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Zubr1.0-VL-4B | I1-IQ3_XXS | 4.4B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Huihui-Qwen3-VL-4B-Instruct-abliterated | I1-IQ3_XXS | 4.4B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-VL-4B-Instruct-Uncensored | I1-IQ3_XXS | 4.4B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| OpenCaption-4B-VL-SFT-v1.0 | I1-IQ3_XXS | 4.4B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Parable-Qwen3-4B-Claude-Fable-5 | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Jan-v1-4B | IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-4b-Z-Image-Turbo-AbliteratedV1 | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-4B-abliterated | IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Neuron-4B-Instruct | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| ChineseErrorCorrector4-4B | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| FastContext-1.0-4B-SFT-abliterated | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-4B-Instruct_NSFW-V2.1 | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| FastContext-1.0-4B-SFT | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| fable-traces-abliterated | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Nexa-AI-4B-Instruct | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Lumen-4B-Instruct | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-4B-Instruct-2507-heretic | IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-HereticLM-4B | I1-IQ3_XXS | 4.0B | 1.56 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-VL-4B-Instruct-Uncensored-abliterated | Q2_K | 4.4B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| PopiT-Qwen3-4B-Medical-SFT-1128 | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Logics-Parsing-v2 | Q2_K | 4.4B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Z-Image-Engineer-V6 | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Huihui-Qwen3-4B-Instruct-2507-abliterated | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Josiefied-Qwen3-4B-abliterated-v2 | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-4B-abliterated-v2 | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| CyberSecQwen-4B | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Huihui-Qwen3-4B-abliterated-v2 | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-Reranker-4B | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Octen-Embedding-4B | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3-Embedding-4B | Q2_K | 4.0B | 1.55 GiB | 5.06 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Qwen3.5-9B | Q4_K_M | 9.7B | 5.47 GiB | 1.13 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| Ministral-8B-Instruct-2410 | Q4_K_S | 8.0B | 4.36 GiB | 2.23 GiB | 7.43 GiB | 0.01 GiB | 48±12.9% |
| nomic-embed-code | Q5_K_S | 7.1B | 4.60 GiB | 1.97 GiB | 7.42 GiB | 0.02 GiB | 48±12.9% |
| SuperGemma-4-12b-abliterated | I1-Q2_K_S | 12.0B | 4.19 GiB | 2.38 GiB | 7.42 GiB | 0.02 GiB | 48±12.9% |
| gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-heretic | I1-Q2_K_S | 12.0B | 4.19 GiB | 2.38 GiB | 7.42 GiB | 0.02 GiB | 48±12.9% |
| gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic | I1-Q2_K_S | 12.0B | 4.19 GiB | 2.38 GiB | 7.42 GiB | 0.02 GiB | 48±12.9% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Measured on this card
| Workload◍ | Median | Middle 50% | Runs |
|---|---|---|---|
| Image generation | 7.02 it/s | 5.47–8.41 | 287 |
Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.
Questions people ask
- What AI models can a GeForce RTX 2070 SUPER run?
- 919 of 2118 indexed open-weight models fit a GeForce RTX 2070 SUPER at 131,072 context with q4_0 KV cache, the largest being OLMoE-1B-7B-0924-Instruct at I1-IQ2_M. That covers text, vision-language, image, video and speech models.
- How much usable memory does a GeForce RTX 2070 SUPER actually have?
- Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a GeForce RTX 2070 SUPER fast for local AI?
- Its memory bandwidth is 448 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.