NVIDIA · consumer

GeForce RTX 5090 D

GeForce RTX 5090 D has 32 GB of VRAM at 1792 GB/s — about 29.76 GiB usable after driver and compositor overhead. 1949 of 2118 indexed models fit at 64K context with q8_0 KV.

Spec sheet· bandwidth, theoreticalFrom the file· fit from summed bytesPredicted· speed
Memory
32 GB
GDDR7
Bandwidth
1792 GB/s
512-bit bus
Tensor FP16
419 TF
dense
TDP
575 W
$2299 MSRP
KV cachef16q8_0q4_0quantizing the KV cache is a ~2× lever on the dominant term at long context
audio tts 21text 1668vision language 177image 2video 16embedding 26audio asr 39

What fits at 64K context

largest quantization that fits, per model · 1949 of 2118 indexed
ModelBest quantParamsWeightsKVTotal in memoryHeadroomtok/s
Qwen3-TTS-12Hz-0.6B-BaseF32915M28.88 GiB0.00 GiB29.72 GiB0.04 GiB44±12.9%
Open_Gpt4_8x7B_v0.1MoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
dolphin-2.6-mixtral-8x7bMoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
dolphin-2.5-mixtral-8x7bMoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
dolphin-2.7-mixtral-8x7bMoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
Nous-Hermes-2-Mixtral-8x7B-DPOMoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
Mixtral-8x7B-v0.1MoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
Mixtral-8x7B-Instruct-v0.1MoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
Mixtral-8x7B-MoE-RP-StoryMoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
Noromaid-v0.4-Mixtral-Instruct-8x7b-ZlossMoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
Open_Gpt4_8x7B_v0.2MoEQ4_046.7B24.63 GiB4.25 GiB29.71 GiB0.05 GiB62±37%
Darwin-35B-A3B-OpusMoEQ6_K_L36.0B28.22 GiB0.66 GiB29.69 GiB0.07 GiB201±37%
Aurora-Code-1MoEQ6_K_L34.7B28.22 GiB0.66 GiB29.69 GiB0.07 GiB201±37%
grug-35b-v2MoEQ6_K_L35.1B28.22 GiB0.66 GiB29.69 GiB0.07 GiB201±37%
grug-35bMoEQ6_K_L35.1B28.22 GiB0.66 GiB29.69 GiB0.07 GiB201±37%
WorldSim-Opus-3.6-35B-A3BMoEQ6_K_L35.1B28.22 GiB0.66 GiB29.69 GiB0.07 GiB201±37%
Qwen3.6-35B-A3B-AnkoMoEQ6_K_L35.1B28.22 GiB0.66 GiB29.69 GiB0.07 GiB201±37%
KAT-Coder-V2.5-DevMoEQ6_K_L34.7B28.22 GiB0.66 GiB29.69 GiB0.07 GiB201±37%
Ornith-1.0-35BMoEQ6_K_L34.7B28.22 GiB0.66 GiB29.69 GiB0.07 GiB201±37%
Nex-N2-miniMoEQ6_K_L35.1B28.22 GiB0.66 GiB29.69 GiB0.07 GiB201±37%
grug-27bQ8_027.4B26.70 GiB2.13 GiB29.68 GiB0.08 GiB44±12.9%
Carnice-V2-27bQ8_027.4B26.70 GiB2.13 GiB29.68 GiB0.08 GiB44±12.9%
Fara1.5-27BQ8_027.4B26.70 GiB2.13 GiB29.68 GiB0.08 GiB44±12.9%
Seed-OSS-36B-Instruct-biprojected-norm-preserving-abliteratedI1-Q4_K_M36.2B20.27 GiB8.50 GiB29.67 GiB0.09 GiB44±12.9%
Seed-OSS-36B-InstructQ4_K_M36.2B20.27 GiB8.50 GiB29.67 GiB0.09 GiB44±12.9%
Hermes-4.3-36B-hereticI1-Q4_K_M36.2B20.27 GiB8.50 GiB29.67 GiB0.09 GiB44±12.9%
Hermes-4.3-36BQ4_K_M36.2B20.27 GiB8.50 GiB29.67 GiB0.09 GiB44±12.9%
Seed-OSS-36B-BaseQ4_K_M36.2B20.27 GiB8.50 GiB29.67 GiB0.09 GiB44±12.9%
DeepSeek-R1-Distill-Llama-70BUD-IQ2_XXS70.6B18.11 GiB10.63 GiB29.66 GiB0.10 GiB44±12.9%
Hermes-4-70BUD-IQ2_XXS70.6B18.10 GiB10.63 GiB29.65 GiB0.11 GiB44±12.9%
Llama-3.3-70B-InstructUD-IQ2_XXS70.6B18.10 GiB10.63 GiB29.65 GiB0.11 GiB44±12.9%
Janus-Pro-7BF167.4B12.88 GiB15.94 GiB29.64 GiB0.12 GiB44±12.9%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTPIQ3_M27.8B26.65 GiB2.13 GiB29.64 GiB0.12 GiB44±12.9%
Qwen3.5-27B-HERETIC-Polaris-Advanced-Thinking-Alpha-uncensoredQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-ThinkingQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.6-27B-Heretic2-ThinkingQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.6-27B-abliteratedQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Webcoda-AI-27BQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
KoQweopus-3.5-27B-experimentalQ8_027.8B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Huihui-Qwen3.6-27B-abliteratedQ8_027.8B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.5-27B-hereticQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Huihui-Qwen3.5-27B-abliteratedQ8_027.8B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.5-Queen-27BQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.5-27B-abliteratedQ8_026.9B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
ThinkingCap-Qwen3.6-27B-hereticQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
MusaCoder-27BQ8_026.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen-Image-BenchQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.5-27B-uncensored-heretic-v1Q8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Bonsai-27B-unpackedQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Ternary-Bonsai-27B-unpackedQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-hereticQ8_027.4B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Darwin-28B-REASONQ8_026.9B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliteratedQ8_027.8B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.5-27B-Claude-4.6-Opus-Reasoning-DistilledQ8_027.8B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.5-27BQ8_027.8B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Qwen3.5-27B-WebNovel-Writer-zhQ8_026.9B26.63 GiB2.13 GiB29.62 GiB0.14 GiB44±12.9%
Assistant_Pepe_70BIQ1_S70.6B18.07 GiB10.63 GiB29.62 GiB0.14 GiB44±12.9%
granite-4.1-30bQ5_128.9B20.19 GiB8.50 GiB29.60 GiB0.16 GiB44±12.9%
Gemma-4-Novelist-Eclipse-31BQ5_K_L32.7B22.77 GiB5.94 GiB29.58 GiB0.18 GiB44±12.9%
Gemma-4-31B-StyleTuneQ5_K_L32.7B22.77 GiB5.94 GiB29.58 GiB0.18 GiB44±12.9%
From the filePredictedwhat these mean

Speed is modeled, not measured: decode is memory-bandwidth bound, so tokens per second is bytes read per token against achievable bandwidth. Mixture-of-experts models carry a wider band because only the routed experts are read each step, and few have been measured publicly.

Measured on this card

third-party benchmarks, aggregated
WorkloadMedianMiddle 50%Runs
Image generation33.31 it/s24.9638.2224
Benchmarked· n=24

Aggregated from community-submitted runs, so the spread is wide by nature — it covers different models, resolutions, step counts and settings, not one controlled configuration. Read the middle 50% rather than the median alone. These figures are reproduced with attribution from vladmandic-sd-data-benchmark, which publishes no licence — so we display and link rather than redistribute them.

Questions people ask

What AI models can a GeForce RTX 5090 D run?
1949 of 2118 indexed open-weight models fit a GeForce RTX 5090 D at 65,536 context with q8_0 KV cache, the largest being Qwen3-TTS-12Hz-0.6B-Base at F32. That covers text, vision-language, image, video and speech models.
How much usable memory does a GeForce RTX 5090 D actually have?
Its nameplate is 32 GB, but about 29.76 GiB is available to a model once driver and compositor overhead is accounted for.
Is a GeForce RTX 5090 D fast for local AI?
Its memory bandwidth is 1792 GB/s, and that figure — not teraflops — is what governs token generation speed. Capacity decides what you can run; bandwidth decides how fast it runs.