AMD · consumer

Radeon RX 6750 GRE

Radeon RX 6750 GRE has 10 GB of VRAM at 320 GB/s — about 9.30 GiB usable after driver and compositor overhead. 1196 of 2118 indexed models fit at 64K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
320 GB/s
160-bit bus
Tensor FP16
dense
TDP
170 W
$269 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1000vision language 99embedding 26audio asr 38video 12audio tts 20image 1

What fits at 64K context

largest quantization that fits, per model · 1196 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3-16B-A3BMoEUD-IQ2_XXS16.0B5.22 GiB3.19 GiB9.30 GiB0.00 GiB28±37%
MiMo-VL-7B-RLI1-Q3_K_M8.3B3.59 GiB4.78 GiB9.30 GiB0.00 GiB24±26.5%
Kuwutu-7B-CYOA-v2I1-Q3_K_M7.6B3.59 GiB4.78 GiB9.30 GiB0.00 GiB24±26.5%
Parable-Granite-4.1-8B-Claude-Fable-5I1-Q2_K8.4B3.05 GiB5.31 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
Tess-4-9BQ6_K9.7B7.30 GiB1.06 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16-abliteratedI1-IQ1_S13.9B3.03 GiB5.31 GiB9.30 GiB0.00 GiB24±26.5%
Ministral-3-14B-Reasoning-2512-UncensoredI1-IQ1_S13.9B3.03 GiB5.31 GiB9.30 GiB0.00 GiB24±26.5%
granite-4.1-8bUD-IQ2_M8.8B3.05 GiB5.31 GiB9.29 GiB0.01 GiB24±26.5%
INTELLECT-1-InstructI1-IQ2_XXS10.2B2.77 GiB5.58 GiB9.29 GiB0.01 GiB24±26.5%
Hy-MT2-7BQ4_K_S8.0B4.09 GiB4.25 GiB9.28 GiB0.02 GiB24±26.5%
Hunyuan-7B-InstructQ4_K_S7.5B4.09 GiB4.25 GiB9.28 GiB0.02 GiB24±26.5%
Smilodon-9B-v1I1-IQ1_M10.2B2.37 GiB5.97 GiB9.28 GiB0.02 GiB24±26.5%
bella-bartender-v2I1-IQ1_M9.2B2.37 GiB5.97 GiB9.28 GiB0.02 GiB24±26.5%
Gemma-The-Writer-9B-HERETIC-Uncensored-AbliteratedI1-IQ1_M9.2B2.37 GiB5.97 GiB9.28 GiB0.02 GiB24±26.5%
Dirty-Muse-Writer-v01-Uncensored-Erotica-NSFWI1-IQ1_M9.2B2.37 GiB5.97 GiB9.28 GiB0.02 GiB24±26.5%
Gemma-2-9B-It-SPPO-Iter3I1-IQ1_M9.2B2.37 GiB5.97 GiB9.28 GiB0.02 GiB24±26.5%
Gemma-SEA-LION-v3-9B-ITI1-IQ1_M9.2B2.37 GiB5.97 GiB9.28 GiB0.02 GiB24±26.5%
G2-Darkest-Writer-Dirty-Shirley-9B-v2I1-IQ1_M9.2B2.37 GiB5.97 GiB9.28 GiB0.02 GiB24±26.5%
G2-Darkest-Writer-9B-v1I1-IQ1_M9.2B2.37 GiB5.97 GiB9.28 GiB0.02 GiB24±26.5%
Tiger-Gemma-9B-v3I1-IQ1_M9.2B2.37 GiB5.97 GiB9.28 GiB0.02 GiB24±26.5%
Ministral-8B-Instruct-2410Q5_K_M8.0B5.33 GiB3.02 GiB9.28 GiB0.02 GiB24±26.5%
NVIDIA-Nemotron-3-Nano-4B-BF16Q4_K_M4.0B2.77 GiB5.58 GiB9.28 GiB0.02 GiB24±26.5%
granite-3.3-8b-instructUD-IQ3_XXS8.2B3.02 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Forsaken-Void-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Silver-Siren-ST-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Tess-3-Mistral-Nemo-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
KrakenSakura-Maelstrom-12B-v1IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Dans-PersonalityEngine-V1.3.0-12bI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
MN-12B-Runeweaver-RP-RUI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Impish_Bloodmoon_12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Wayfarer-2-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Wayfarer-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Muse-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
MN-Violet-Lotus-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Rocinante-X-12B-v1-Heretic-UncensoredI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Mistral-Heretica-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
arcee-fusion-lumaid-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Mistral-NeMo-12B-AbliteratedI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Captain-Eris_Violet-V0.420-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Rocinante-X-12B-v1-absolute-heresyI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Rocinante-X-12B-v1I1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Mistral-Nemo-2407-12B-Thinking-Claude-Gemini-GPT5.2-Uncensored-HERETICI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Dans-SakuraKaze-V1.0.0-12bI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Mistral-Nemo-Inst-2407-12B-Thinking-Uncensored-HERETIC-HI-Claude-OpusI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Mistral-Nemo-Instruct-2407-12B-Thinking-M-Claude-Opus-High-ReasoningI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Mordant-12B-ThinkI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Riverfish-Rocinante-12B-SFT-DPOI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
MN-Violet-Lotus-12B-HereticI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Himeyuri-Magnum-12B-HereticMergeI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Kinggaroo-12b-v1I1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Magnum-Picaro-0.7-v2-12bI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
magnum-v2.5-12b-ktoI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Mistral-Nemo-Prism-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Nera_Noctis-12BI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
Chronos-Gold-12B-1.0I1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
patricide-12B-Unslop-MellI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
MN-12B-Celeste-V1.9I1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
magnum-v2-12bI1-IQ1_M12.2B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation2.78 it/s2.772.925
Benchmarked· n=5

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6750 GRE run?
1196 of 2118 indexed open-weight models fit a Radeon RX 6750 GRE at 65,536 context with q8_0 KV cache, the largest being Qwen3-16B-A3B at UD-IQ2_XXS. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6750 GRE actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6750 GRE fast for local AI?
Its memory bandwidth is 320 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.