NVIDIA · consumer

GeForce RTX 3060 OEM

GeForce RTX 3060 OEM has 6 GB of VRAM at 336 GB/s — about 5.58 GiB usable after driver and compositor overhead. 1230 of 2118 indexed models fit at 4K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
336 GB/s
192-bit bus
Tensor FP16
57 TF
dense
TDP
185 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1052vision language 91audio asr 38embedding 26image 1audio tts 19video 3

What fits at 4K context

largest quantization that fits, per model · 1230 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Nous-Hermes-2-SOLAR-10.7BQ3_K_S10.7B4.34 GiB0.40 GiB5.58 GiB0.00 GiB50±12.9%
SOLAR-10.7B-Instruct-v1.0I1-Q3_K_S10.7B4.34 GiB0.40 GiB5.58 GiB0.00 GiB50±12.9%
Luna-7B-A4BMoEI1-Q5_K_M6.7B4.47 GiB0.30 GiB5.58 GiB0.00 GiB51±37%
Falcon3-10B-InstructIQ3_M10.3B4.38 GiB0.33 GiB5.58 GiB0.00 GiB50±12.9%
Jan-v2-VL-highQ4_08.8B4.45 GiB0.30 GiB5.58 GiB0.00 GiB50±12.9%
Jan-v2-VL-medQ4_08.8B4.45 GiB0.30 GiB5.58 GiB0.00 GiB50±12.9%
ReasonCritic-7BQ4_08.2B4.45 GiB0.30 GiB5.58 GiB0.00 GiB50±12.9%
mythos-9b-unhingedQ4_08.2B4.45 GiB0.30 GiB5.58 GiB0.00 GiB50±12.9%
Qwen3-8B-BaseQ4_08.2B4.45 GiB0.30 GiB5.58 GiB0.00 GiB50±12.9%
MiniCPM-o-4_5Q4_09.4B4.45 GiB0.30 GiB5.58 GiB0.00 GiB50±12.9%
Hy-MT2-7BQ4_18.0B4.47 GiB0.27 GiB5.58 GiB0.00 GiB50±12.9%
deepseek-llm-7b-chatQ4_K_S6.9B3.75 GiB1.00 GiB5.58 GiB0.00 GiB50±12.9%
Smilodon-9B-v1I1-IQ3_S10.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
bella-bartender-v2I1-IQ3_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-IQ3_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-IQ3_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
Gemma-2-9B-It-SPPO-Iter3I1-IQ3_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
Gemma-SEA-LION-v3-9B-ITI1-IQ3_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-IQ3_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
G2-Darkest-Writer-9B-v1I1-IQ3_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
Tiger-Gemma-9B-v3I1-IQ3_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
gemma-2-9b-it-abliteratedQ3_K_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
gemma-2-9b-itQ3_K_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
Tiger-Gemma-9B-v1Q3_K_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
magnum-v4-9bQ3_K_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
gemma-2-9bQ3_K_S9.2B4.04 GiB0.70 GiB5.58 GiB0.00 GiB50±12.9%
Marco-Nano-InstructMoEI1-Q4_K_S8.0B4.57 GiB0.23 GiB5.58 GiB0.00 GiB185±37%
Trinity-Nano-PreviewMoEQ6_K6.1B4.72 GiB0.08 GiB5.58 GiB0.00 GiB188±37%
gemma-4-E4B-it-hereticQ4_18.0B4.69 GiB0.07 GiB5.58 GiB0.00 GiB50±12.9%
Hunyuan-7B-InstructQ4_17.5B4.47 GiB0.27 GiB5.57 GiB0.01 GiB50±12.9%
DeepSeek-R1-Distill-Llama-8B-AbliteratedI1-IQ2_XXS8.0B4.47 GiB0.27 GiB5.57 GiB0.01 GiB50±12.9%
gemma-4-E4B-uncensoredI1-IQ4_XS7.9B4.69 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
gemma-4-E4B-it-qat-q4_0-unquantized-hereticI1-IQ4_XS7.9B4.69 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
gemma-4-E4B-it-qat-heretic_decensoredI1-IQ4_XS7.9B4.69 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
gemma-4-E4B-it-QAT-SOMPOA-heresyI1-IQ4_XS7.9B4.69 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
gemma4-e4b-mahou-nsfwI1-IQ4_XS7.9B4.69 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
gemma-4-E4B-it-mentalchat16kI1-IQ4_XS7.9B4.69 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
gemma4-E4B-it-abliteratedI1-IQ4_XS7.9B4.69 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
gemma-4-E4B-it-OBLITERATEDI1-IQ4_XS8.0B4.69 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
gemma-3n-E4B-itQ5_K_M7.8B4.68 GiB0.06 GiB5.57 GiB0.01 GiB50±12.9%
canary-qwen-2.5bBF162.6B4.73 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
EXAONE-Deep-7.8BQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
EXAONE-3.5-7.8B-InstructQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
deepseek-math-7b-instructQ4_K_S6.9B3.75 GiB1.00 GiB5.57 GiB0.01 GiB50±12.9%
nomic-embed-codeQ5_K_S7.1B4.60 GiB0.12 GiB5.57 GiB0.01 GiB50±12.9%
Janus-Pro-7BI1-Q4_K_S7.4B3.75 GiB1.00 GiB5.57 GiB0.01 GiB50±12.9%
deepseek-coder-7b-instruct-v1.5I1-Q4_K_S6.9B3.75 GiB1.00 GiB5.57 GiB0.01 GiB50±12.9%
Gemma-4-E4B-LuchadorQ3_K_L8.0B4.69 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
Qwen3-TTS-12Hz-0.6B-BaseQ4_K_M915M4.72 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
granite-3.1-8b-instructQ4_K_S8.2B4.41 GiB0.33 GiB5.57 GiB0.01 GiB50±12.9%
VoxCPM2F162.3B4.72 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
Ministral-3-8B-Instruct-2512-BF16Q3_K_L8.9B4.45 GiB0.28 GiB5.57 GiB0.01 GiB50±12.9%
OLMo-2-1124-7B-InstructQ3_K_L7.3B3.68 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
Olmo-3-7B-InstructQ3_K_L7.3B3.68 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
Olmo-3-7B-ThinkI1-Q3_K_L7.3B3.68 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
ERNIE-21B-A3B-Thinking-Gemini-3-Pro-High-Reasoning-V2I1-IQ1_M21.8B4.63 GiB0.12 GiB5.56 GiB0.02 GiB50±12.9%
ERNIE-21B-A3B-Claude-4.5-High-OPUS-ThinkingI1-IQ1_M21.8B4.63 GiB0.12 GiB5.56 GiB0.02 GiB50±12.9%
ERNIE-4.5-21B-A3B-ThinkingI1-IQ1_M21.8B4.63 GiB0.12 GiB5.56 GiB0.02 GiB50±12.9%
t5-v1_1-xxlQ2_K4.8B4.72 GiB0.00 GiB5.56 GiB0.02 GiB50±12.9%
NousCoder-14BIQ2_XS14.8B4.37 GiB0.33 GiB5.56 GiB0.02 GiB51±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a GeForce RTX 3060 OEM run?
1230 of 2118 indexed open-weight models fit a GeForce RTX 3060 OEM at 4,096 context with q8_0 KV cache, the largest being Nous-Hermes-2-SOLAR-10.7B at Q3_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3060 OEM actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3060 OEM fast for local AI?
Its memory bandwidth is 336 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.