NVIDIA · consumer
GeForce RTX 4070 GDDR6
GeForce RTX 4070 GDDR6 has 12 GB of VRAM at 480 GB/s — about 11.16 GiB usable after driver and compositor overhead. 876 of 2118 indexed models fit at 128K context with q8_0 KV.
Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
12 GB
GDDR6
Bandwidth
480 GB/s
192-bit bus
Tensor FP16
117 TF
dense
TDP
200 W
$599 MSRP
video 14vision language 95audio tts 20text 687audio asr 38embedding 21image 1
What fits at 128K context
largest quantization that fits, per model · 876 of 2118 indexed
| Model | Best quant | Params | Weights● | KV● | Total in memory◐ | Headroom◐ | tok/s◐ |
|---|---|---|---|---|---|---|---|
| Wan2.1-FLF2V-14B-720P | Q4_1 | 16.4B | 10.32 GiB | 0.00 GiB | 11.16 GiB | 0.00 GiB | 33±12.9% |
| Wan2.1-I2V-14B-480P | Q4_1 | 16.4B | 10.32 GiB | 0.00 GiB | 11.15 GiB | 0.01 GiB | 33±12.9% |
| Wan2.1-I2V-14B-720P | Q4_1 | 16.4B | 10.32 GiB | 0.00 GiB | 11.15 GiB | 0.01 GiB | 33±12.9% |
| Muse-Glimmer-30B | IQ2_S | 29.8B | 9.35 GiB | 0.91 GiB | 11.15 GiB | 0.01 GiB | 34±12.9% |
| orpheus-3b-0.1-pretrained | Q6_K | 3.8B | 2.90 GiB | 7.44 GiB | 11.15 GiB | 0.01 GiB | 33±12.9% |
| DeepSeek-Coder-V2-Lite-BaseMoE | I1-Q4_0 | 15.7B | 8.32 GiB | 2.02 GiB | 11.14 GiB | 0.02 GiB | 59±37% |
| Voxtral-Mini-4B-Realtime-2602KV unresolved | Q6_K | 4.4B | 3.41 GiB | 6.91 GiB | 11.13 GiB | 0.03 GiB | 33±12.9% |
| DeepSeek-Coder-V2-Lite-InstructMoE | IQ4_NL | 15.7B | 8.29 GiB | 2.02 GiB | 11.12 GiB | 0.04 GiB | 59±37% |
| DeepSeek-V2-Lite-ChatMoE | IQ4_NL | 15.7B | 8.29 GiB | 2.02 GiB | 11.12 GiB | 0.04 GiB | 59±37% |
| Ministral-3-3B-Instruct-2512 | Q8_0 | 3.8B | 3.40 GiB | 6.91 GiB | 11.12 GiB | 0.04 GiB | 33±12.9% |
| Ministral-3-3B-Reasoning-2512 | Q8_0 | 4.3B | 3.40 GiB | 6.91 GiB | 11.12 GiB | 0.04 GiB | 33±12.9% |
| Ministral-3-3B-Instruct-2512-BF16 | Q8_0 | 4.3B | 3.40 GiB | 6.91 GiB | 11.12 GiB | 0.04 GiB | 33±12.9% |
| Amaretto-3B | Q8_0 | 4.3B | 3.40 GiB | 6.91 GiB | 11.12 GiB | 0.04 GiB | 33±12.9% |
| ERNIE-4.5-21B-A3B-PT | UD-IQ1_S | 21.9B | 6.58 GiB | 3.72 GiB | 11.12 GiB | 0.04 GiB | 33±12.9% |
| Marco-Nano-InstructMoE | I1-IQ2_S | 8.0B | 2.91 GiB | 7.44 GiB | 11.12 GiB | 0.04 GiB | 26±37% |
| Wan2.2-Distill-Models | Q5_1 | 14.3B | 10.27 GiB | 0.00 GiB | 11.10 GiB | 0.06 GiB | 33±12.9% |
| Bernini-R | Q5_1 | 14.3B | 10.26 GiB | 0.00 GiB | 11.10 GiB | 0.06 GiB | 34±12.9% |
| SkyReels-V2-DF-14B-540P | Q5_1 | 14.3B | 10.27 GiB | 0.00 GiB | 11.10 GiB | 0.06 GiB | 34±12.9% |
| Gemma-4-12B-StyleTune | I1-IQ3_M | 13.0B | 5.74 GiB | 4.50 GiB | 11.09 GiB | 0.07 GiB | 34±12.9% |
| gemma-4-12b-heretic-styletune-head | I1-IQ3_M | 12.0B | 5.74 GiB | 4.50 GiB | 11.09 GiB | 0.07 GiB | 34±12.9% |
| syrian-gemma-12b | I1-IQ3_M | 13.0B | 5.74 GiB | 4.50 GiB | 11.09 GiB | 0.07 GiB | 34±12.9% |
| InternVL3_5-14B | Q5_K_L | 15.1B | 10.24 GiB | 0.00 GiB | 11.08 GiB | 0.08 GiB | 34±12.9% |
| Phi-4-mini-reasoning | Q3_K_S | 3.8B | 1.77 GiB | 8.50 GiB | 11.08 GiB | 0.08 GiB | 34±12.9% |
| Phi-4-mini-instruct | Q3_K_S | 3.8B | 1.77 GiB | 8.50 GiB | 11.08 GiB | 0.08 GiB | 34±12.9% |
| Yi-1.5-6B-Chat | Q8_0 | 6.1B | 6.00 GiB | 4.25 GiB | 11.07 GiB | 0.09 GiB | 34±12.9% |
| Trinity-MiniMoE | Q2_K_L | 26.1B | 9.15 GiB | 1.12 GiB | 11.07 GiB | 0.09 GiB | 85±37% |
| ERNIE-4.5-21B-A3B-Thinking | UD-IQ1_S | 21.8B | 6.53 GiB | 3.72 GiB | 11.07 GiB | 0.09 GiB | 34±12.9% |
| Aura-4B | I1-IQ3_XXS | 4.5B | 1.75 GiB | 8.50 GiB | 11.06 GiB | 0.10 GiB | 34±12.9% |
| magnum-v2-4b | I1-IQ3_XXS | 4.5B | 1.75 GiB | 8.50 GiB | 11.06 GiB | 0.10 GiB | 34±12.9% |
| Impish_LLAMA_4B | IQ3_XXS | 4.5B | 1.75 GiB | 8.50 GiB | 11.06 GiB | 0.10 GiB | 34±12.9% |
| granite-20b-code-instruct-8k | IQ4_XS | 20.1B | 10.19 GiB | 0.00 GiB | 11.06 GiB | 0.10 GiB | 34±12.9% |
| granite-20b-code-base-8k | I1-IQ4_XS | 20.1B | 10.19 GiB | 0.00 GiB | 11.06 GiB | 0.10 GiB | 34±12.9% |
| Kimi-VL-A3B-Thinking-2506MoE | IQ4_XS | 16.4B | 8.21 GiB | 2.02 GiB | 11.05 GiB | 0.11 GiB | 59±37% |
| Kimi-VL-A3B-InstructMoE | IQ4_XS | 16.4B | 8.21 GiB | 2.02 GiB | 11.05 GiB | 0.11 GiB | 59±37% |
| Qwen3.6-28BMoE | I1-IQ2_M | 28.2B | 8.91 GiB | 1.33 GiB | 11.04 GiB | 0.12 GiB | 86±37% |
| Qwen3.5-28BMoE | I1-IQ2_M | 28.7B | 8.91 GiB | 1.33 GiB | 11.04 GiB | 0.12 GiB | 86±37% |
| GLM-4.7-Flash-REAP-23B-A3BMoE | UD-IQ1_S | 23.0B | 6.71 GiB | 3.51 GiB | 11.03 GiB | 0.13 GiB | 44±37% |
| OLMoE-1B-7B-0924-InstructMoE | I1-IQ2_XXS | 6.9B | 1.76 GiB | 8.50 GiB | 11.03 GiB | 0.13 GiB | 22±37% |
| Llama-3.1-Minitron-4B-Width-Base | Q2_K | 4.5B | 1.71 GiB | 8.50 GiB | 11.03 GiB | 0.13 GiB | 34±12.9% |
| Qwen3.6-35B-A3B-REAM-160-ru-agentMoE | IQ3_XXS | 23.6B | 8.89 GiB | 1.33 GiB | 11.02 GiB | 0.14 GiB | 83±37% |
| Tiger-Gemma-12B-v3 | IQ3_M | 12.8B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| AfriqueGemma-12B | I1-IQ3_M | 12.2B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| SuperGemma-4-12b-abliterated | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-heretic | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12B-coder-fable5-composer2.5-v1-uncensored-heretic | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12B-it-uncensored-heretic | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12B-it-qat-q4_0-unquantized-uncensored-heretic | Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| Grug-12B | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| Aura-Medium-v1-BF16 | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12B-it-Esper4 | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12B-it-Guardpoint | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| Gemma-4-12B-it-AEON-Abliterated-K4-BF16 | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12B-it-Tachibana-Agent | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12b-marvin-gutenberg-rp-v2 | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12b-crownelius-writer | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| Huihui-gemma-4-12B-coder-fable5-composer2.5-v1-abliterated | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma-4-12b-asterion-agentic | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| Huihui-gemma-4-12B-agentic-fable5-abliterated | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| g4-12b-it-trismegistus | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
| gemma4-12b-it-asimov | I1-Q3_K_M | 12.0B | 5.67 GiB | 4.50 GiB | 11.01 GiB | 0.15 GiB | 34±12.9% |
Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.
Questions people ask
- What AI models can a GeForce RTX 4070 GDDR6 run?
- 876 of 2118 indexed open-weight models fit a GeForce RTX 4070 GDDR6 at 131,072 context with q8_0 KV cache, the largest being Wan2.1-FLF2V-14B-720P at Q4_1. That covers text, vision-language, image, video and speech models.
- How much usable memory does a GeForce RTX 4070 GDDR6 actually have?
- Its nameplate is 12 GB, but about 11.16 GiB is available to a model once driver and compositor overhead is accounted for.
- Is a GeForce RTX 4070 GDDR6 fast for local AI?
- Its memory bandwidth is 480 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.