AMD · consumer

Radeon RX 6750 GRE

Radeon RX 6750 GRE has 10 GB of VRAM at 320 GB/s — about 9.30 GiB usable after driver and compositor overhead. 1527 of 2118 indexed models fit at 16K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
320 GB/s
160-bit bus
Tensor FP16
dense
TDP
170 W
$269 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1313vision language 114video 12audio asr 39image 2audio tts 21embedding 26

What fits at 16K context

largest quantization that fits, per model · 1527 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
L3-8B-Sunfall-v0.5-Stheno-v3.2Q6_K_L8.0B7.30 GiB1.06 GiB9.30 GiB0.00 GiB24±26.5%
Qwen3-48B-A4B-Savant-Commander-Distill-12X-Closed-Open-Heretic-UncensoredMoEI1-IQ1_M33.6B7.19 GiB1.20 GiB9.30 GiB0.00 GiB41±37%
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
Le-Chaton-Slim-23BMoEI1-Q2_K_S23.3B7.52 GiB0.86 GiB9.30 GiB0.00 GiB41±37%
Qwen3-VL-30B-A3B-InstructMoEUD-TQ1_031.1B7.60 GiB0.80 GiB9.29 GiB0.01 GiB63±37%
Ministral-3-8B-Instruct-2512-BF16Q6_K_L8.9B7.22 GiB1.13 GiB9.29 GiB0.01 GiB24±26.5%
GLM-4.7-Flash-REAP-23B-A3B-absolute-heresyMoEI1-Q2_K23.0B7.94 GiB0.44 GiB9.29 GiB0.01 GiB69±37%
EuroLLM-22B-Instruct-2512IQ2_XS22.6B6.53 GiB1.79 GiB9.28 GiB0.02 GiB24±26.5%
Qwen3-VL-30B-A3B-ThinkingMoEUD-TQ1_031.1B7.59 GiB0.80 GiB9.28 GiB0.02 GiB63±37%
Qwen3-30B-A3B-Thinking-2507MoEUD-TQ1_030.5B7.59 GiB0.80 GiB9.28 GiB0.02 GiB63±37%
medgemma-27b-itUD-IQ2_XXS28.8B7.31 GiB0.99 GiB9.28 GiB0.02 GiB24±26.5%
gemma-3-27b-itUD-IQ2_XXS27.4B7.31 GiB0.99 GiB9.28 GiB0.02 GiB24±26.5%
medgemma-27b-text-itUD-IQ2_XXS27.0B7.31 GiB0.99 GiB9.28 GiB0.02 GiB24±26.5%
Llama-3.2-11B-Vision-InstructQ4_K_S10.7B7.01 GiB1.33 GiB9.28 GiB0.02 GiB24±26.5%
Falcon3-7B-InstructQ8_07.5B7.38 GiB0.93 GiB9.28 GiB0.02 GiB24±26.5%
Qwen3-30B-A3BMoEIQ2_XXS30.5B7.59 GiB0.80 GiB9.28 GiB0.02 GiB63±37%
Pantheon-Proto-RP-1.8-30B-A3BMoEIQ2_XXS30.5B7.59 GiB0.80 GiB9.28 GiB0.02 GiB63±37%
granite-3.1-2b-instructQ8_02.5B7.70 GiB0.66 GiB9.27 GiB0.03 GiB24±26.5%
codegeex4-all-9bIQ1_M9.4B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
ERNIE-21B-A3B-Thinking-Gemini-3-Pro-High-Reasoning-V2I1-IQ3_XXS21.8B7.88 GiB0.46 GiB9.26 GiB0.04 GiB24±26.5%
ERNIE-21B-A3B-Claude-4.5-High-OPUS-ThinkingI1-IQ3_XXS21.8B7.88 GiB0.46 GiB9.26 GiB0.04 GiB24±26.5%
ERNIE-4.5-21B-A3B-ThinkingI1-IQ3_XXS21.8B7.88 GiB0.46 GiB9.26 GiB0.04 GiB24±26.5%
glm-4-9b-chatIQ1_M9.4B3.00 GiB5.31 GiB9.26 GiB0.04 GiB24±26.5%
LFM2-8B-A1BMoEQ8_08.3B8.26 GiB0.10 GiB9.25 GiB0.05 GiB68±37%
gemma-2-9bQ5_19.2B6.52 GiB1.79 GiB9.25 GiB0.05 GiB24±26.5%
gemma-2-9b-itQ5_19.2B6.52 GiB1.79 GiB9.25 GiB0.05 GiB24±26.5%
GLM-4.6V-FlashQ6_K_L10.3B7.98 GiB0.33 GiB9.25 GiB0.05 GiB24±26.5%
GLM-Z1-9B-0414Q6_K_L9.4B7.98 GiB0.33 GiB9.25 GiB0.05 GiB24±26.5%
GLM-4-9B-0414Q6_K_L9.4B7.98 GiB0.33 GiB9.25 GiB0.05 GiB24±26.5%
FrickFritz-4BF164.7B8.07 GiB0.27 GiB9.25 GiB0.05 GiB24±26.5%
Newton-bot-3-VLM-mini-4BF164.7B8.07 GiB0.27 GiB9.25 GiB0.05 GiB24±26.5%
qwen3.5-4b-agentic-coder-v4F164.7B8.07 GiB0.27 GiB9.25 GiB0.05 GiB24±26.5%
Myth-4BF164.3B8.07 GiB0.27 GiB9.25 GiB0.05 GiB24±26.5%
Qwen3.5-4B-UncensoredF164.7B8.07 GiB0.27 GiB9.25 GiB0.05 GiB24±26.5%
JOSIE-2-4B-PreviewF164.7B8.07 GiB0.27 GiB9.25 GiB0.05 GiB24±26.5%
Qwen3.5-4BBF164.7B8.07 GiB0.27 GiB9.25 GiB0.05 GiB24±26.5%
Surogate-3.5-4BF165.3B8.07 GiB0.27 GiB9.25 GiB0.05 GiB24±26.5%
Qwopus3.5-4B-v3BF164.7B8.07 GiB0.27 GiB9.25 GiB0.05 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16-abliteratedIQ4_XS13.9B6.96 GiB1.33 GiB9.25 GiB0.05 GiB24±26.5%
Ministral-3-14B-abliteratedIQ4_XS13.9B6.96 GiB1.33 GiB9.25 GiB0.05 GiB24±26.5%
Ministral-3-14B-Reasoning-2512-UncensoredIQ4_XS13.9B6.96 GiB1.33 GiB9.25 GiB0.05 GiB24±26.5%
gemma-3-12b-itQ4_012.2B7.52 GiB0.78 GiB9.25 GiB0.05 GiB24±26.5%
Forsaken-Void-12BI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Silver-Siren-ST-12BI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Tess-3-Mistral-Nemo-12BI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
KrakenSakura-Maelstrom-12B-v1Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
MN-12B-Runeweaver-RP-RUI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Impish_Bloodmoon_12BI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Wayfarer-2-12BI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Wayfarer-12BI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Muse-12BI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Vikhr-Nemo-12B-Instruct-R-21-09-24Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Mistral-Nemo-Base-2407Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
writing-roleplay-20k-context-nemo-12b-v1.0Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Dans-PersonalityEngine-V1.3.0-12bI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
mini-magnum-12b-v1.1Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Lumimaid-v0.2-12BQ4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
MN-Violet-Lotus-12BI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
Rocinante-X-12B-v1-Heretic-UncensoredI1-Q4_K_M12.2B6.96 GiB1.33 GiB9.24 GiB0.06 GiB24±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation2.78 it/s2.772.925
Benchmarked· n=5

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6750 GRE run?
1527 of 2118 indexed open-weight models fit a Radeon RX 6750 GRE at 16,384 context with q8_0 KV cache, the largest being L3-8B-Sunfall-v0.5-Stheno-v3.2 at Q6_K_L. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6750 GRE actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6750 GRE fast for local AI?
Its memory bandwidth is 320 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.