NVIDIA · consumer

GeForce RTX 2070

GeForce RTX 2070 has 8 GB of VRAM at 448 GB/s — about 7.44 GiB usable after driver and compositor overhead. 1440 of 2118 indexed models fit at 4K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
8 GB
GDDR6
Bandwidth
448 GB/s
256-bit bus
Tensor FP16
60 TF
dense
TDP
175 W
$499 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1242vision language 103video 8audio tts 21embedding 26image 2audio asr 38

What fits at 4K context

largest quantization that fits, per model · 1440 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
OmniAtlas-Qwen3-30B-A3BI1-IQ1_M31.7B6.59 GiB0.00 GiB7.44 GiB0.00 GiB48±12.9%
Qwen3-Omni-30B-A3B-CaptionerI1-IQ1_M31.7B6.59 GiB0.00 GiB7.44 GiB0.00 GiB48±12.9%
QwenPaw-Flash-9BQ6_K_S9.4B6.56 GiB0.04 GiB7.44 GiB0.00 GiB48±12.9%
HomunculusIQ4_XS12.5B6.41 GiB0.18 GiB7.43 GiB0.01 GiB48±12.9%
NVIDIA-Nemotron-Nano-12B-v2IQ4_XS12.3B6.29 GiB0.27 GiB7.43 GiB0.01 GiB48±12.9%
Nexa-AI-4x4B-InstructMoEI1-Q4_012.1B6.46 GiB0.16 GiB7.43 GiB0.01 GiB50±37%
glm-4-9b-chat-1mQ4_K_M9.5B5.88 GiB0.70 GiB7.42 GiB0.02 GiB48±12.9%
Aurora-Code-1MoEI1-IQ1_M34.7B6.59 GiB0.02 GiB7.42 GiB0.02 GiB250±37%
Parable-Granite-4.1-8B-Claude-Fable-5I1-Q6_K8.4B6.41 GiB0.18 GiB7.42 GiB0.02 GiB48±12.9%
Le-Chaton-Slim-23BMoEI1-IQ2_XS23.3B6.49 GiB0.11 GiB7.41 GiB0.03 GiB101±37%
LFM2.5-8B-A1BMoEUD-Q6_K8.5B6.60 GiB0.01 GiB7.41 GiB0.03 GiB142±37%
glm-4v-9bQ5_K_M13.9B6.57 GiB0.00 GiB7.41 GiB0.03 GiB48±12.9%
Apertus-8B-Instruct-2509Q6_K_L8.1B6.40 GiB0.14 GiB7.41 GiB0.03 GiB49±12.9%
ERNIE-4.5-21B-A3B-ThinkingUD-IQ1_S21.8B6.53 GiB0.06 GiB7.41 GiB0.03 GiB48±12.9%
NVIDIA-Nemotron-Nano-9B-v2Q5_K_S8.9B6.32 GiB0.25 GiB7.41 GiB0.03 GiB48±12.9%
openNemo-9B-abliteratedQ5_K_S8.9B6.32 GiB0.25 GiB7.41 GiB0.03 GiB48±12.9%
OLMo-2-1124-13B-InstructQ3_K_S13.7B5.68 GiB0.88 GiB7.41 GiB0.03 GiB48±12.9%
Smilodon-9B-v1I1-Q5_K_M10.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
bella-bartender-v2I1-Q5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-Q5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-Q5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
Gemma-2-9B-It-SPPO-Iter3I1-Q5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
Gemma-SEA-LION-v3-9B-ITI1-Q5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-Q5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
G2-Darkest-Writer-9B-v1I1-Q5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
Tiger-Gemma-9B-v3I1-Q5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
gemma-2-9b-it-abliteratedQ5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
gemma-2-9b-itQ5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
Tiger-Gemma-9B-v1Q5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
magnum-v4-9bQ5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
gemma-2-9bQ5_K_M9.2B6.19 GiB0.37 GiB7.40 GiB0.04 GiB48±12.9%
Gemma-3-27B-MeditronFOI1-IQ1_S28.8B6.26 GiB0.26 GiB7.40 GiB0.04 GiB49±12.9%
gemma-4-E4B-it-hereticQ6_K8.0B6.55 GiB0.03 GiB7.40 GiB0.04 GiB48±12.9%
Nemotron-Mini-4B-InstructQ2_K4.2B6.43 GiB0.14 GiB7.39 GiB0.05 GiB48±12.9%
granite-4.0-7B-A1B-Creative-v0.1MoEQ8_06.7B6.62 GiB0.01 GiB7.39 GiB0.05 GiB164±37%
InternVL3_5-8BQ6_K_L8.5B6.54 GiB0.00 GiB7.39 GiB0.05 GiB48±12.9%
HunyuanVideo-1.5Q6_K8.3B6.54 GiB0.00 GiB7.39 GiB0.05 GiB49±12.9%
Devstral-Small-2-24B-Instruct-2512UD-IQ2_XXS24.0B6.29 GiB0.18 GiB7.38 GiB0.06 GiB49±12.9%
Mistral-Small-3.2-24B-Instruct-2506UD-IQ2_XXS24.0B6.29 GiB0.18 GiB7.38 GiB0.06 GiB49±12.9%
Devstral-Small-2507UD-IQ2_XXS23.6B6.29 GiB0.18 GiB7.38 GiB0.06 GiB49±12.9%
Devstral-Small-2505UD-IQ2_XXS23.6B6.29 GiB0.18 GiB7.38 GiB0.06 GiB49±12.9%
Magistral-Small-2509UD-IQ2_XXS24.0B6.29 GiB0.18 GiB7.38 GiB0.06 GiB49±12.9%
Magistral-Small-2507UD-IQ2_XXS23.6B6.29 GiB0.18 GiB7.38 GiB0.06 GiB49±12.9%
Mistral-Small-3.1-24B-Instruct-2503UD-IQ2_XXS24.0B6.29 GiB0.18 GiB7.38 GiB0.06 GiB49±12.9%
Magistral-Small-2506UD-IQ2_XXS23.6B6.29 GiB0.18 GiB7.38 GiB0.06 GiB49±12.9%
Tess-4-9BQ5_K_M9.7B6.51 GiB0.04 GiB7.38 GiB0.06 GiB48±12.9%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopQ4_K_S12.1B6.49 GiB0.03 GiB7.37 GiB0.07 GiB49±12.9%
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresyMoEI1-IQ2_S23.0B6.51 GiB0.06 GiB7.37 GiB0.07 GiB171±37%
solar-pro-preview-instructKV unresolvedIQ2_XS22.1B6.16 GiB0.35 GiB7.37 GiB0.07 GiB49±12.9%
codegeex4-all-9bQ4_K_M9.4B5.82 GiB0.70 GiB7.37 GiB0.07 GiB49±12.9%
glm-4-9b-chat-abliteratedQ4_K_M9.4B5.82 GiB0.70 GiB7.37 GiB0.07 GiB49±12.9%
glm-4-9b-chatQ4_K_M9.4B5.82 GiB0.70 GiB7.37 GiB0.07 GiB49±12.9%
Llama3.2-24B-A3B-II-Dark-Champion-INSTRUCT-Heretic-Abliterated-UncensoredMoEI1-Q2_K18.0B6.44 GiB0.12 GiB7.37 GiB0.07 GiB134±37%
Ministral-8B-Instruct-2410Q6_K_L8.0B6.38 GiB0.16 GiB7.37 GiB0.07 GiB49±12.9%
Assistant_Pepe_8BQ6_K_L6.39 GiB0.14 GiB7.37 GiB0.07 GiB49±12.9%
dolphin-2.9.2-Phi-3-MediumKV unresolvedQ3_K14.0B6.29 GiB0.22 GiB7.37 GiB0.07 GiB49±12.9%
v6-Finch-14B-HFQ2_K_L14.1B5.45 GiB1.07 GiB7.37 GiB0.07 GiB49±12.9%
Grug-12BIQ4_XS12.0B6.32 GiB0.20 GiB7.36 GiB0.08 GiB49±12.9%
gemma-4-12B-it-Esper4IQ4_XS12.0B6.32 GiB0.20 GiB7.36 GiB0.08 GiB49±12.9%
gemma-4-12B-itIQ4_XS12.0B6.32 GiB0.20 GiB7.36 GiB0.08 GiB49±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation5.75 it/s4.026.72293
Benchmarked· n=293

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 2070 run?
1440 of 2118 indexed open-weight models fit a GeForce RTX 2070 at 4,096 context with q4_0 KV cache, the largest being OmniAtlas-Qwen3-30B-A3B at I1-IQ1_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 2070 actually have?
Its nameplate is 8 GB, but about 7.44 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 2070 fast for local AI?
Its memory bandwidth is 448 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.