AMD · consumer

Radeon RX 6750 GRE

Radeon RX 6750 GRE has 10 GB of VRAM at 320 GB/s — about 9.30 GiB usable after driver and compositor overhead. 1415 of 2118 indexed models fit at 64K context with q4_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
320 GB/s
160-bit bus
Tensor FP16
dense
TDP
170 W
$269 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
text 1207vision language 110embedding 26video 12audio tts 21audio asr 38image 1

What fits at 64K context

largest quantization that fits, per model · 1415 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Falcon3-10B-InstructQ4_010.3B5.52 GiB2.81 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
Fimbulvetr-11B-v2I1-Q3_K_M10.7B4.98 GiB3.38 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3.6-28BMoEI1-IQ2_XS28.2B8.04 GiB0.35 GiB9.29 GiB0.01 GiB91±37%
Qwen3.5-28BMoEI1-IQ2_XS28.7B8.04 GiB0.35 GiB9.29 GiB0.01 GiB91±37%
stable-code-3bQ8_02.8B2.77 GiB5.63 GiB9.29 GiB0.01 GiB24±26.5%
rocket-3BQ8_02.8B2.77 GiB5.63 GiB9.29 GiB0.01 GiB24±26.5%
phi-2Q8_02.8B2.75 GiB5.63 GiB9.29 GiB0.01 GiB24±26.5%
DeepSeek-R1-Distill-Llama-8B-AbliteratedI1-IQ3_XXS8.0B6.10 GiB2.25 GiB9.29 GiB0.01 GiB24±26.5%
MiMo-VL-7B-RLI1-Q6_K8.3B5.83 GiB2.53 GiB9.29 GiB0.01 GiB24±26.5%
Kuwutu-7B-CYOA-v2I1-Q6_K7.6B5.83 GiB2.53 GiB9.29 GiB0.01 GiB24±26.5%
Ministral-3-8B-Instruct-2512-BF16Q5_K_L8.9B5.95 GiB2.39 GiB9.28 GiB0.02 GiB24±26.5%
Qwen3.6-35B-A3B-REAM-160-ru-agentMoEQ2_K_S23.6B8.02 GiB0.35 GiB9.28 GiB0.02 GiB86±37%
deepseek-coder-1.3b-baseF321.3B5.02 GiB3.38 GiB9.28 GiB0.02 GiB24±26.5%
OpenCaption-2B-VL-SFT-v1.0F322.1B6.42 GiB1.97 GiB9.28 GiB0.02 GiB24±26.5%
gemma-4-12B-it-qat-q4_0-unquantized-uncensored-hereticQ4_012.0B7.07 GiB1.26 GiB9.28 GiB0.02 GiB24±26.5%
Gemma-4-12B-StyleTuneI1-Q4_K_S13.0B7.07 GiB1.26 GiB9.27 GiB0.03 GiB24±26.5%
gemma-4-12b-heretic-styletune-headI1-Q4_K_S12.0B7.07 GiB1.26 GiB9.27 GiB0.03 GiB24±26.5%
syrian-gemma-12bI1-Q4_K_S13.0B7.07 GiB1.26 GiB9.27 GiB0.03 GiB24±26.5%
MiroThinker-v1.0-8BQ5_K_L8.2B5.81 GiB2.53 GiB9.27 GiB0.03 GiB24±26.5%
Qwen3-8B-abliteratedQ5_K_L8.2B5.81 GiB2.53 GiB9.27 GiB0.03 GiB24±26.5%
Qwen3-8BQ5_K_L8.2B5.81 GiB2.53 GiB9.27 GiB0.03 GiB24±26.5%
Josiefied-Qwen3-8B-abliterated-v1Q5_K_L8.2B5.81 GiB2.53 GiB9.27 GiB0.03 GiB24±26.5%
Nemotron-Orchestrator-8BQ5_K_L8.2B5.81 GiB2.53 GiB9.27 GiB0.03 GiB24±26.5%
DeepSeek-R1-0528-Qwen3-8BQ5_K_L8.2B5.81 GiB2.53 GiB9.27 GiB0.03 GiB24±26.5%
Ministral-3-14B-Instruct-2512-BF16Q2_K_L13.9B5.50 GiB2.81 GiB9.26 GiB0.04 GiB24±26.5%
Phi-3-medium-128k-instructQ2_K14.0B4.79 GiB3.52 GiB9.26 GiB0.04 GiB24±26.5%
Phi-3-medium-4k-instructI1-Q2_K14.0B4.79 GiB3.52 GiB9.26 GiB0.04 GiB24±26.5%
Qwen3-16B-A3BMoEQ3_K_S16.0B6.68 GiB1.69 GiB9.26 GiB0.04 GiB39±37%
GLM-4.7-Flash-DerestrictedMoEI1-IQ2_XXS31.2B7.42 GiB0.93 GiB9.26 GiB0.04 GiB60±37%
Huihui-GLM-4.7-Flash-abliteratedMoEI1-IQ2_XXS31.2B7.42 GiB0.93 GiB9.26 GiB0.04 GiB60±37%
Qwythos-9B-v2Q6_K_L9.7B7.76 GiB0.56 GiB9.26 GiB0.04 GiB24±26.5%
Tess-4-9BQ6_K_L9.7B7.76 GiB0.56 GiB9.26 GiB0.04 GiB24±26.5%
Qwen3-VL-Embedding-8BQ6_K8.1B5.79 GiB2.53 GiB9.25 GiB0.05 GiB24±26.5%
Qwen3-8B-BaseQ6_K8.2B5.79 GiB2.53 GiB9.25 GiB0.05 GiB24±26.5%
qwen-indic-v1I1-Q6_K7.6B5.79 GiB2.53 GiB9.25 GiB0.05 GiB24±26.5%
Qwen3-Embedding-8BQ6_K7.6B5.79 GiB2.53 GiB9.25 GiB0.05 GiB24±26.5%
gemma-3-12b-it-vl-Gemini-3-Pro-Preview-Heretic-Uncensored-ThinkingI1-Q4_112.2B7.04 GiB1.26 GiB9.24 GiB0.06 GiB24±26.5%
gemma-3-12b-it-vl-Deepseek-v3.1-Heretic-Uncensored-ThinkingI1-Q4_112.2B7.04 GiB1.26 GiB9.24 GiB0.06 GiB24±26.5%
gemma-3-12b-it-vl-GLM-4.7-Flash-Heretic-Uncensored-ThinkingI1-Q4_112.2B7.04 GiB1.26 GiB9.24 GiB0.06 GiB24±26.5%
Floppa-12B-Gemma3-UncensoredI1-Q4_112.2B7.04 GiB1.26 GiB9.24 GiB0.06 GiB24±26.5%
gemma-3-12b-it-hereticI1-Q4_112.2B7.04 GiB1.26 GiB9.24 GiB0.06 GiB24±26.5%
gemma-3-12b-it-abliteratedQ4_112.2B7.04 GiB1.26 GiB9.24 GiB0.06 GiB24±26.5%
gemma-3-12b-itQ4_112.2B7.04 GiB1.26 GiB9.24 GiB0.06 GiB24±26.5%
Phi-4-reasoning-plusIQ2_M14.7B4.76 GiB3.52 GiB9.24 GiB0.06 GiB24±26.5%
Phi-4-reasoningIQ2_M14.7B4.76 GiB3.52 GiB9.24 GiB0.06 GiB24±26.5%
phi-4IQ2_M14.7B4.76 GiB3.52 GiB9.24 GiB0.06 GiB24±26.5%
Maestro1-9BQ5_18.8B5.77 GiB2.53 GiB9.23 GiB0.07 GiB24±26.5%
Jan-v2-VL-highQ5_18.8B5.77 GiB2.53 GiB9.23 GiB0.07 GiB24±26.5%
Jan-v2-VL-medQ5_18.8B5.77 GiB2.53 GiB9.23 GiB0.07 GiB24±26.5%
Olmo-3-7B-InstructQ6_K7.3B5.58 GiB2.72 GiB9.23 GiB0.07 GiB24±26.5%
Olmo-3-7B-ThinkI1-Q6_K7.3B5.58 GiB2.72 GiB9.23 GiB0.07 GiB24±26.5%
MiniCPM-o-4_5Q5_19.4B5.77 GiB2.53 GiB9.23 GiB0.07 GiB24±26.5%
GLM-Z1-32B-0414UD-IQ1_S32.6B7.17 GiB1.07 GiB9.23 GiB0.07 GiB24±26.5%
GLM-4-32B-0414UD-IQ1_S32.6B7.17 GiB1.07 GiB9.23 GiB0.07 GiB24±26.5%
Qwen3.6-12B-IQ-Ultra-Heretic-Uncensored-Thinking-V2-HightopQ5_K_M12.1B7.84 GiB0.42 GiB9.23 GiB0.07 GiB24±26.5%
Qwen3.5-27B-Engineer-Deckard-GeminiI1-IQ2_XXS27.7B7.14 GiB1.13 GiB9.23 GiB0.07 GiB24±26.5%
Qwen3.5-27B-HERETIC-Polaris-Advanced-Thinking-Alpha-uncensoredI1-IQ2_XXS27.4B7.14 GiB1.13 GiB9.23 GiB0.07 GiB24±26.5%
Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-ThinkingI1-IQ2_XXS27.4B7.14 GiB1.13 GiB9.23 GiB0.07 GiB24±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation2.78 it/s2.772.925
Benchmarked· n=5

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6750 GRE run?
1415 of 2118 indexed open-weight models fit a Radeon RX 6750 GRE at 65,536 context with q4_0 KV cache, the largest being Falcon3-10B-Instruct at Q4_0. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6750 GRE actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6750 GRE fast for local AI?
Its memory bandwidth is 320 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.