AMD · consumer

Radeon RX 6750 GRE

Radeon RX 6750 GRE has 10 GB of VRAM at 320 GB/s — about 9.30 GiB usable after driver and compositor overhead. 784 of 2118 indexed models fit at 128K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
10 GB
GDDR6
Bandwidth
320 GB/s
160-bit bus
Tensor FP16
dense
TDP
170 W
$269 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
audio tts 19text 609vision language 86embedding 21video 12audio asr 36image 1

What fits at 128K context

largest quantization that fits, per model · 784 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
orpheus-3b-0.1-ftUD-IQ1_M3.8B0.95 GiB7.44 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-instruct-8kIQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
granite-20b-code-base-8kI1-IQ3_S20.1B8.32 GiB0.00 GiB9.30 GiB0.00 GiB24±26.5%
Llama-Doctor-3.2-3B-InstructI1-IQ2_XXS3.2B0.95 GiB7.44 GiB9.29 GiB0.01 GiB24±26.5%
Llama-3.2-3B-Instruct-roleplay-tunedI1-IQ2_XXS3.2B0.95 GiB7.44 GiB9.29 GiB0.01 GiB24±26.5%
Llama-3.2-3B-Instruct-heretic-ablitered-uncensoredI1-IQ2_XXS3.2B0.95 GiB7.44 GiB9.29 GiB0.01 GiB24±26.5%
Llama3.2-3B-creative-writer-v0.1I1-IQ2_XXS3.2B0.95 GiB7.44 GiB9.29 GiB0.01 GiB24±26.5%
Firefly-V3.2I1-IQ2_XXS3.2B0.95 GiB7.44 GiB9.29 GiB0.01 GiB24±26.5%
Firefly-V3I1-IQ2_XXS3.2B0.95 GiB7.44 GiB9.29 GiB0.01 GiB24±26.5%
Felldude-Uncensored-Ministral3-3B-bf16I1-IQ3_XS3.8B1.47 GiB6.91 GiB9.29 GiB0.01 GiB24±26.5%
Ministral-3-3B-Instruct-2512-BF16IQ3_XS4.3B1.47 GiB6.91 GiB9.29 GiB0.01 GiB24±26.5%
Amaretto-3BI1-IQ3_XS4.3B1.47 GiB6.91 GiB9.29 GiB0.01 GiB24±26.5%
Qwen3-1.7BIQ3_M2.0B0.96 GiB7.44 GiB9.29 GiB0.01 GiB24±26.5%
OpenClaude-1.7B-MergedIQ4_XS1.7B0.96 GiB7.44 GiB9.29 GiB0.01 GiB24±26.5%
Ministral-8B-Instruct-2410IQ4_XS8.0B4.14 GiB4.21 GiB9.29 GiB0.01 GiB24±26.5%
Ling-mini-2.0MoEQ2_K_L16.3B5.74 GiB2.66 GiB9.28 GiB0.02 GiB36±37%
Qwen3-VL-Reranker-2BIQ4_XS2.1B0.95 GiB7.44 GiB9.28 GiB0.02 GiB24±26.5%
Atomight-V2.5-1.7BIQ4_XS1.7B0.95 GiB7.44 GiB9.28 GiB0.02 GiB24±26.5%
OpenCaption-2B-VL-SFT-v1.0IQ4_XS2.1B0.95 GiB7.44 GiB9.28 GiB0.02 GiB24±26.5%
gaon-1.7b-v2-instructIQ4_XS1.7B0.95 GiB7.44 GiB9.28 GiB0.02 GiB24±26.5%
gaon-1.7b-v2-translateIQ4_XS1.7B0.95 GiB7.44 GiB9.28 GiB0.02 GiB24±26.5%
Llama-3.2-3B-Instruct-abliteratedI1-IQ1_S3.6B0.93 GiB7.44 GiB9.27 GiB0.03 GiB24±26.5%
nomic-embed-codeQ5_K_S7.1B4.60 GiB3.72 GiB9.27 GiB0.03 GiB24±26.5%
Supertron2-Reranker-2BI1-IQ4_XS2.1B0.94 GiB7.44 GiB9.27 GiB0.03 GiB24±26.5%
Uni-MuMER-Qwen3-VL-2BI1-IQ4_XS2.1B0.94 GiB7.44 GiB9.27 GiB0.03 GiB24±26.5%
Qwen3-VL-2B-ThinkingIQ4_XS2.1B0.94 GiB7.44 GiB9.27 GiB0.03 GiB24±26.5%
Qwen3-VL-2B-InstructIQ4_XS2.1B0.94 GiB7.44 GiB9.27 GiB0.03 GiB24±26.5%
Lightning-1.7BIQ4_XS1.7B0.94 GiB7.44 GiB9.27 GiB0.03 GiB24±26.5%
DorsetHeatwaveLLM2I1-IQ4_XS1.7B0.94 GiB7.44 GiB9.27 GiB0.03 GiB24±26.5%
Fara1.5-9BQ4_K_L9.4B6.21 GiB2.13 GiB9.27 GiB0.03 GiB24±26.5%
QwenPaw-Flash-9BQ4_K_L9.4B6.21 GiB2.13 GiB9.27 GiB0.03 GiB24±26.5%
grug-9bQ4_K_L9.4B6.21 GiB2.13 GiB9.27 GiB0.03 GiB24±26.5%
OmniCoder-9BQ4_K_L9.4B6.21 GiB2.13 GiB9.27 GiB0.03 GiB24±26.5%
Ornith-1.0-9BQ4_K_L9.2B6.21 GiB2.13 GiB9.27 GiB0.03 GiB24±26.5%
Qwen3.5-9B-NeoQ4_K_L9.7B6.21 GiB2.13 GiB9.27 GiB0.03 GiB24±26.5%
ERNIE-21B-A3B-Thinking-Gemini-3-Pro-High-Reasoning-V2I1-IQ1_M21.8B4.63 GiB3.72 GiB9.27 GiB0.03 GiB24±26.5%
ERNIE-21B-A3B-Claude-4.5-High-OPUS-ThinkingI1-IQ1_M21.8B4.63 GiB3.72 GiB9.27 GiB0.03 GiB24±26.5%
ERNIE-4.5-21B-A3B-ThinkingI1-IQ1_M21.8B4.63 GiB3.72 GiB9.27 GiB0.03 GiB24±26.5%
Qwythos-9B-v2Q5_K_S9.7B6.21 GiB2.13 GiB9.27 GiB0.03 GiB24±26.5%
Tess-4-9BQ5_K_S9.7B6.21 GiB2.13 GiB9.27 GiB0.03 GiB24±26.5%
Qwen3-VL-Embedding-2BQ3_K_L2.1B0.93 GiB7.44 GiB9.27 GiB0.03 GiB24±26.5%
Qwen3.6-35B-A3B-REAM-160-ru-agentMoEIQ2_XS23.6B7.03 GiB1.33 GiB9.26 GiB0.04 GiB53±37%
Qwen3.5-9B-CoderI1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Qwopus3.5-9B-v3.5Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Qwythos-9B-Claude-Mythos-5-1M-MTPI1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Huihui-Qwythos-9B-Claude-Mythos-5-1M-abliteratedI1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Qwen3.5-9B-Fable-5-v1I1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
PINQWEN-3.5-9B-1M-BF16I1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Openprose-2-FlashI1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Qwen3.5-9B-Nikusui-v1I1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Ornstein-3.5-9B-V1.5I1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Ornith-1.0-9B-heretic-MTPI1-Q5_K_M9.4B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Qwythos-9B-Claude-Mythos-5-1MQ5_K9.4B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
dotwebs-1I1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
liftQ5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Hemlock-Qwopus3.5-9B-CoderI1-Q5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Qwen3.5-9B-DeepSeek-V4-FlashQ5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
Qwen3.5-9BQ5_K_M9.7B6.19 GiB2.13 GiB9.25 GiB0.05 GiB24±26.5%
SuperGemma-4-12b-abliteratedI1-IQ2_S12.0B3.80 GiB4.50 GiB9.24 GiB0.06 GiB24±26.5%
gemma-4-12B-agentic-fable5-composer2.5-v2-3.5x-tau2-uncensored-hereticI1-IQ2_S12.0B3.80 GiB4.50 GiB9.24 GiB0.06 GiB24±26.5%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation2.78 it/s2.772.925
Benchmarked· n=5

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a Radeon RX 6750 GRE run?
784 of 2118 indexed open-weight models fit a Radeon RX 6750 GRE at 131,072 context with q8_0 KV cache, the largest being orpheus-3b-0.1-ft at UD-IQ1_M. That covers text, vision-language, image, video and speech models.
How much usable memory does a Radeon RX 6750 GRE actually have?
Its nameplate is 10 GB, but about 9.30 GiB is available to a model once driver and compositor overhead is accounted for.
Is a Radeon RX 6750 GRE fast for local AI?
Its memory bandwidth is 320 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.