NVIDIA · consumer

GeForce RTX 3060 OEM

GeForce RTX 3060 OEM has 6 GB of VRAM at 336 GB/s — about 5.58 GiB usable after driver and compositor overhead. 1115 of 2118 indexed models fit at 16K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
6 GB
GDDR6
Bandwidth
336 GB/s
192-bit bus
Tensor FP16
57 TF
dense
TDP
185 W
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
vision language 88text 941audio asr 38audio tts 19embedding 26video 3

What fits at 16K context

largest quantization that fits, per model · 1115 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwythos-9B-v2Q3_K_S9.7B4.48 GiB0.27 GiB5.58 GiB0.00 GiB50±12.9%
Tess-4-9BQ3_K_S9.7B4.48 GiB0.27 GiB5.58 GiB0.00 GiB50±12.9%
dolphin-2.9.3-mistral-7B-32kIQ4_XS7.2B3.68 GiB1.06 GiB5.58 GiB0.00 GiB50±12.9%
Mistral-7B-Instruct-v0.3-ParasiteIQ4_XS7.2B3.68 GiB1.06 GiB5.58 GiB0.00 GiB50±12.9%
Mistral-7B-Instruct-v0.3-JbliteratedIQ4_XS7.2B3.68 GiB1.06 GiB5.58 GiB0.00 GiB50±12.9%
Mistral-7B-v0.3IQ4_XS7.2B3.68 GiB1.06 GiB5.58 GiB0.00 GiB50±12.9%
OpenChat-3.5-7B-Qwen-v2.0KV unresolvedIQ4_XS7.2B3.67 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
Mistral-7B-v0.2IQ4_XS7.2B3.67 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
ContextualKunoichi_KTO-7BIQ4_XS7.2B3.67 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
mistral-7b-uncensoredKV unresolvedIQ4_XS7.2B3.67 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
Yarn-Mistral-7b-128kKV unresolvedIQ4_XS7.2B3.67 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
Ninja-v1-RP-WIPKV unresolvedIQ4_XS7.2B3.67 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
Silicon-Maid-7BKV unresolvedIQ4_XS7.2B3.67 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
SpydazWeb_AI_CyberTron_Ultra_7bKV unresolvedIQ4_XS7.2B3.67 GiB1.06 GiB5.57 GiB0.01 GiB50±12.9%
AMALIA-9B-0626-DPOQ2_K9.2B3.35 GiB1.39 GiB5.57 GiB0.01 GiB50±12.9%
canary-qwen-2.5bBF162.6B4.73 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
EXAONE-Deep-7.8BQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
EXAONE-3.5-7.8B-InstructQ4_K_L7.8B4.73 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_baseI1-Q4_08.1B4.50 GiB0.23 GiB5.57 GiB0.01 GiB50±12.9%
Ministral-3-8B-Instruct-2512-BF16-abliteratedI1-Q3_K_S8.9B3.60 GiB1.13 GiB5.57 GiB0.01 GiB50±12.9%
Amaretto-8BI1-Q3_K_S8.9B3.60 GiB1.13 GiB5.57 GiB0.01 GiB50±12.9%
Ministral-3-8B-Instruct-2512Q3_K_S8.9B3.60 GiB1.13 GiB5.57 GiB0.01 GiB50±12.9%
Ministral-3-8B-Reasoning-2512Q3_K_S8.9B3.60 GiB1.13 GiB5.57 GiB0.01 GiB50±12.9%
Qwen3-TTS-12Hz-0.6B-BaseQ4_K_M915M4.72 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
VoxCPM2F162.3B4.72 GiB0.00 GiB5.57 GiB0.01 GiB50±12.9%
GLM-4.6V-FlashIQ3_M10.3B4.40 GiB0.33 GiB5.57 GiB0.01 GiB50±12.9%
glm4.1v-9b-base-sftI1-IQ3_M10.3B4.40 GiB0.33 GiB5.57 GiB0.01 GiB50±12.9%
GLM-Z1-9B-0414IQ3_M9.4B4.40 GiB0.33 GiB5.57 GiB0.01 GiB50±12.9%
GLM-4-9B-0414IQ3_M9.4B4.40 GiB0.33 GiB5.57 GiB0.01 GiB50±12.9%
gemma-4-E2B-itQ8_05.1B4.70 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
gemma-4-E2B-itQ8_05.1B4.70 GiB0.07 GiB5.57 GiB0.01 GiB50±12.9%
LFM2.5-8B-A1BMoEUD-Q4_K_S8.5B4.67 GiB0.10 GiB5.57 GiB0.01 GiB136±37%
t5-v1_1-xxlQ2_K4.8B4.72 GiB0.00 GiB5.56 GiB0.02 GiB50±12.9%
Phi-3.5-mini-instructIQ3_S3.8B1.57 GiB3.19 GiB5.56 GiB0.02 GiB50±12.9%
NuExtract-1.5Q3_K_S3.8B1.57 GiB3.19 GiB5.56 GiB0.02 GiB50±12.9%
Phi-3.5-mini-instructQ3_K_S3.8B1.57 GiB3.19 GiB5.56 GiB0.02 GiB50±12.9%
Phi-3.5-mini-instruct_UncensoredIQ3_S3.8B1.57 GiB3.19 GiB5.56 GiB0.02 GiB50±12.9%
Phi-3-mini-128k-instructQ3_K_S3.8B1.57 GiB3.19 GiB5.56 GiB0.02 GiB50±12.9%
Phi-3-mini-4k-instructQ3_K_S3.8B1.57 GiB3.19 GiB5.56 GiB0.02 GiB50±12.9%
octo-netQ3_K_S3.8B1.57 GiB3.19 GiB5.56 GiB0.02 GiB50±12.9%
granite-3.1-2b-instructQ6_K2.5B4.10 GiB0.66 GiB5.56 GiB0.02 GiB50±12.9%
Ministral-8B-Instruct-2410IQ3_M8.0B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
gemma-4-12BIQ2_S12.0B3.93 GiB0.78 GiB5.56 GiB0.02 GiB50±12.9%
Jan-v3-4B-base-instructQ6_K_L4.4B3.55 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
next-8bI1-IQ3_S8.2B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Supertron2-Reranker-8BI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
next-ocrI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Qwen3-VL-8B-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Midas-FableAgent-8BI1-IQ3_S8.2B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Qwen3-VL-8B-Heretic-1.3.0I1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Qwen3-VL-8B-Thinking-Unredacted-MAXI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Qwen3-VL-8B-Instruct-Minecraft-MT-en-zhI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Qwen-3-VL-8B-Instruct-hereticI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Poe-8B-GLM5-Opus4.6-Sonnet4.5-Kimi-Grok-Gemini-3-pro-preview-HERETICI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
ToolCUA-8BI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Huihui-Qwen3-VL-8B-Instruct-abliteratedI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Qwen3-VL-Reranker-8BI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Salience-1-9BI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Qwen3-VL-8B-Instruct-Uncensored-V2I1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
Maestro1-9BI1-IQ3_S8.8B3.53 GiB1.20 GiB5.56 GiB0.02 GiB50±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Questions people ask

What AI models can a GeForce RTX 3060 OEM run?
1115 of 2118 indexed open-weight models fit a GeForce RTX 3060 OEM at 16,384 context with q8_0 KV cache, the largest being Qwythos-9B-v2 at Q3_K_S. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 3060 OEM actually have?
Its nameplate is 6 GB, but about 5.58 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 3060 OEM fast for local AI?
Its memory bandwidth is 336 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.
GeForce RTX 3060 OEM — what AI models can it run locally? — ossmodeldb