GPU comparison for local inference

GeForce RTX 5090 D vs Titan V

Titan V holds more — 32 GB against 32 GB, which is what decides whether a model runs at all. GeForce RTX 5090 D has 2.06× the memory bandwidth, which is what decides how fast it generates.

Spec sheet· specsFrom the file· fitPredicted· speed

Specifications

GeForce RTX 5090 DTitan V
Memory32 GB32 GB
Usable to GPU32 GB32 GB
Bandwidth1792 GB/s868 GB/s
Memory typeGDDR7HBM2
Bus width512-bit4096-bit
Tensor FP16 (dense)419 TFLOPS119 TFLOPS
TDP575 W250 W
MSRP at launch$2299$5999
Models that fit19541954
Spec sheetPredictedwhat these mean

Same model, both cards

modeled tokens per second at 32K context
ModelQuantParamsGeForce RTX 5090 DTitan VDifference
Qwen3-Coder-30B-A3B-InstructMoEQ6_K30.5B11960+97%
Qwen3.6-27BQ6_K27.8B5427+102%
Qwen3.6-35B-A3BMoEUD-Q6_K36.0B204107+90%
Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTPIQ3_M27.8B4422+103%
Qwen3.8-27BQ6_K27.8B5427+102%
Qwen3.5-9BBF169.7B6934+101%
gemma-4-26B-A4B-itMoEQ8_026.5B4623+103%
gemma-4-12B-itBF1612.0B5125+102%
nemotron-3.5-asr-streaming-0.6bF32638M393224+75%
Qwen3.5-4BBF164.7B13167+96%
gemma-4-12B-it-qat-q4_0-unquantizedQ4_012.0B13268+96%
Qwythos-9B-Claude-Mythos-5-1MQ8_09.4B6633+101%

Token rates are modeled from memory bandwidth, so on models both cards can hold the ratio tracks bandwidth closely. That is the honest shape of the answer: for inference, capacity decides what you can run and bandwidth decides how fast. Neither is teraflops.

How to read this

We earn nothing from either of these cards. If both hold the models you care about, the faster one wins on bandwidth alone. If one holds a model the other cannot, that difference usually matters far more than any percentage of token rate — a model that does not fit runs 5–20× slower once it starts spilling to system memory, not slightly slower.