GPU comparison for local inference

GeForce RTX 5090 D vs GeForce RTX 5080

GeForce RTX 5090 D holds more — 32 GB against 16 GB, which is what decides whether a model runs at all. GeForce RTX 5090 D has 1.87× the memory bandwidth, which is what decides how fast it generates. Of the models indexed here, 266 fit only on GeForce RTX 5090 D and 0 fit only on GeForce RTX 5080.

Spec sheet· specsFrom the file· fitPredicted· speed

Specifications

GeForce RTX 5090 DGeForce RTX 5080
Memory32 GB16 GB
Usable to GPU32 GB16 GB
Bandwidth1792 GB/s960 GB/s
Memory typeGDDR7GDDR7
Bus width512-bit256-bit
Tensor FP16 (dense)419 TFLOPS225 TFLOPS
TDP575 W360 W
MSRP at launch$2299$999
Models that fit19541688
Spec sheetPredictedwhat these mean

Same model, both cards

modeled tokens per second at 32K context
ModelQuantParamsGeForce RTX 5090 DGeForce RTX 5080Difference
Qwen3-Coder-30B-A3B-InstructMoEQ6_K30.5B11988+34%
Qwen3.6-27BQ6_K27.8B5450+9%
Qwen3.6-35B-A3BMoEUD-Q6_K36.0B204192+6%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTPIQ3_M27.8B4449-9%
Qwen3.8-27BQ6_K27.8B5450+9%
Qwen3.5-9BBF169.7B6966+4%
gemma-4-26B-A4B-itMoEQ8_026.5B4649-5%
gemma-4-12B-itBF1612.0B5155-7%
nemotron-3.5-asr-streaming-0.6bF32638M393243+61%
Qwen3.5-4BBF164.7B13174+78%
gemma-4-12B-it-qat-q4_0-unquantizedQ4_012.0B13274+78%
Qwythos-9B-Claude-Mythos-5-1MQ8_09.4B6651+29%

Token rates are modeled from memory bandwidth, so on models both cards can hold the ratio tracks bandwidth closely. That is the honest shape of the answer: for inference, capacity decides what you can run and bandwidth decides how fast. Neither is teraflops.

How to read this

We earn nothing from either of these cards. If both hold the models you care about, the faster one wins on bandwidth alone. If one holds a model the other cannot, that difference usually matters far more than any percentage of token rate — a model that does not fit runs 5–20× slower once it starts spilling to system memory, not slightly slower.