NVIDIA · consumer

GeForce RTX 2070

GeForce RTX 2070 has 8 GB of VRAM at 448 GB/s — about 7.44 GiB usable after driver and compositor overhead. 1407 of 2118 indexed models fit at 8K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
GDDR6
Bandwidth
448 GB/s
256-bit bus
Tensor FP16
60 TF
dense
TDP
175 W
$499 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1210audio tts 21vision language 102image 2video 8embedding 26audio asr 38

What fits at 8K context

largest quantization that fits, per model · 1407 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Gemma-The-Writer-N-Restless-Quill-10B-UncensoredQ4_K_S10.0B5.40 GiB1.19 GiB7.44 GiB0.00 GiB48±12.9%
OmniAtlas-Qwen3-30B-A3BI1-IQ1_M31.7B6.59 GiB0.00 GiB7.44 GiB0.00 GiB48±12.9%
Qwen3-Omni-30B-A3B-CaptionerI1-IQ1_M31.7B6.59 GiB0.00 GiB7.44 GiB0.00 GiB48±12.9%
gemma-7bI1-Q4_K_S8.5B4.70 GiB1.86 GiB7.43 GiB0.01 GiB48±12.9%
Gemma-4-12B-StyleTuneI1-Q3_K_M13.0B6.07 GiB0.51 GiB7.43 GiB0.01 GiB48±12.9%
gemma-4-12b-heretic-styletune-headI1-Q3_K_M12.0B6.07 GiB0.51 GiB7.43 GiB0.01 GiB48±12.9%
syrian-gemma-12bI1-Q3_K_M13.0B6.07 GiB0.51 GiB7.43 GiB0.01 GiB48±12.9%
orpheus-3b-0.1-ftF163.8B6.16 GiB0.46 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-27B-Engineer-Deckard-GeminiI1-IQ1_M27.7B6.30 GiB0.27 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-27B-HERETIC-Polaris-Advanced-Thinking-Alpha-uncensoredI1-IQ1_M27.4B6.30 GiB0.27 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-ThinkingI1-IQ1_M27.4B6.30 GiB0.27 GiB7.43 GiB0.01 GiB48±12.9%
Huihui-Qwen3.5-27B-abliteratedI1-IQ1_M27.8B6.30 GiB0.27 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-27B-Unredacted-MAXI1-IQ1_M27.4B6.30 GiB0.27 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-27B-hereticI1-IQ1_M27.4B6.30 GiB0.27 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-27B-DerestrictedI1-IQ1_M27.8B6.30 GiB0.27 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3.5-27B-Claude-4.6-Opus-Reasoning-DistilledI1-IQ1_M27.8B6.30 GiB0.27 GiB7.43 GiB0.01 GiB48±12.9%
SambaLingo-Japanese-ChatI1-Q5_K_S6.9B4.48 GiB2.13 GiB7.43 GiB0.01 GiB48±12.9%
Qwen3-Coder-REAP-25B-A3BMoEIQ2_XXS24.9B6.24 GiB0.40 GiB7.42 GiB0.02 GiB137±37%
DeepSeek-Coder-V2-Lite-BaseMoEI1-IQ3_XXS15.7B6.49 GiB0.13 GiB7.42 GiB0.02 GiB153±37%
DeepSeek-Coder-V2-Lite-InstructMoEIQ3_XXS15.7B6.49 GiB0.13 GiB7.42 GiB0.02 GiB153±37%
DeepSeek-V2-Lite-ChatMoEIQ3_XXS15.7B6.49 GiB0.13 GiB7.42 GiB0.02 GiB153±37%
GLM-4.7-Flash-DerestrictedMoEI1-IQ1_M31.2B6.39 GiB0.22 GiB7.42 GiB0.02 GiB168±37%
Huihui-GLM-4.7-Flash-abliteratedMoEI1-IQ1_M31.2B6.39 GiB0.22 GiB7.42 GiB0.02 GiB168±37%
v6-Finch-7B-HFQ4_07.6B4.45 GiB2.13 GiB7.42 GiB0.02 GiB48±12.9%
rwkv-6-world-7bQ4_07.6B4.45 GiB2.13 GiB7.42 GiB0.02 GiB48±12.9%
Apriel-1.6-15b-ThinkerI1-IQ3_XS14.9B5.77 GiB0.80 GiB7.42 GiB0.02 GiB48±12.9%
granite-4.0-7B-A1B-Creative-v0.1MoEQ8_06.7B6.62 GiB0.03 GiB7.41 GiB0.03 GiB160±37%
glm-4v-9bQ5_K_M13.9B6.57 GiB0.00 GiB7.41 GiB0.03 GiB48±12.9%
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresyMoEI1-IQ2_XS23.0B6.38 GiB0.22 GiB7.41 GiB0.03 GiB151±37%
deepseek-math-7b-instructQ5_K_M6.9B4.59 GiB1.99 GiB7.41 GiB0.03 GiB48±12.9%
deepseek-llm-7b-chatQ5_K_M6.9B4.59 GiB1.99 GiB7.41 GiB0.03 GiB48±12.9%
Janus-Pro-7BI1-Q5_K_M7.4B4.59 GiB1.99 GiB7.41 GiB0.03 GiB48±12.9%
deepseek-coder-7b-instruct-v1.5I1-Q5_K_M6.9B4.59 GiB1.99 GiB7.41 GiB0.03 GiB48±12.9%
deepseek-coder-6.7b-instructQ5_K_M6.7B4.46 GiB2.13 GiB7.41 GiB0.03 GiB48±12.9%
deepseek-coder-6.7b-baseQ5_K_M6.7B4.46 GiB2.13 GiB7.41 GiB0.03 GiB48±12.9%
deepseek-coder-6.7B-kexerI1-Q5_K_M6.7B4.46 GiB2.13 GiB7.41 GiB0.03 GiB48±12.9%
Magicoder-S-DS-6.7BI1-Q5_K_M6.7B4.46 GiB2.13 GiB7.41 GiB0.03 GiB48±12.9%
MathCoder2-CodeLlama-7BQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
CodeLlama-7b-instruct-hfQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
CodeLlama-7b-hfQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
WizardLM-7B-UncensoredI1-Q5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
Llama-2-7B-32K-InstructI1-Q5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
Luna-AI-Llama2-UncensoredI1-Q5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
Swallow-7b-NVE-instruct-hfI1-Q5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
llava-v1.5-7bQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
CodeLlama-7b-python-hfQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
Wizard-Vicuna-7B-UncensoredQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
llama2_7b_chat_uncensoredQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
WizardLM-7B-V1.0-UncensoredQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
Llama-2-7b-hfQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
pygmalion-2-7bQ5_K_M6.7B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
OLMo-2-1124-7B-InstructQ4_K_L7.3B4.45 GiB2.13 GiB7.40 GiB0.04 GiB48±12.9%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopQ4_012.1B6.43 GiB0.10 GiB7.40 GiB0.04 GiB49±12.9%
gemma-3-12b-it-vl-Gemini-3-Pro-Preview-Heretic-Uncensored-ThinkingI1-Q3_K_L12.2B6.04 GiB0.51 GiB7.40 GiB0.04 GiB48±12.9%
gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-ThinkingI1-Q3_K_L12.2B6.04 GiB0.51 GiB7.40 GiB0.04 GiB48±12.9%
gemma-3-12b-it-ultra-uncensored-hereticQ3_K_L12.2B6.04 GiB0.51 GiB7.40 GiB0.04 GiB48±12.9%
gemma-3-12b-it-vl-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-Q3_K_L12.2B6.04 GiB0.51 GiB7.40 GiB0.04 GiB48±12.9%
Floppa-12B-Gemma3-UncensoredI1-Q3_K_L12.2B6.04 GiB0.51 GiB7.40 GiB0.04 GiB48±12.9%
gemma-3-12b-it-hereticI1-Q3_K_L12.2B6.04 GiB0.51 GiB7.40 GiB0.04 GiB48±12.9%
gemma-3-12b-it-abliteratedQ3_K_L12.2B6.04 GiB0.51 GiB7.40 GiB0.04 GiB48±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation5.75 it/s4.026.72293
Benchmarked· n=293

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 2070 run?
1407 of 2118 indexed open-weight models fit a GeForce RTX 2070 at 8,192 context with q8_0 KV cache, the largest being Gemma-The-Writer-N-Restless-Quill-10B-Uncensored at Q4_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 2070 actually have?
Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 2070 fast for local AI?
Its memory bandwidth is 448 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.