GeForce RTX 5090 vs GeForce RTX 4070 Ti Super
GeForce RTX 5090 holds more — 32 GB against 16 GB, which is what decides whether a model runs at all. GeForce RTX 5090 has 2.67× the memory bandwidth, which is what decides how fast it generates. Of the models indexed here, 266 fit only on GeForce RTX 5090 and 0 fit only on GeForce RTX 4070 Ti Super.
Specifications
| GeForce RTX 5090 | GeForce RTX 4070 Ti Super | |
|---|---|---|
| Memory○ | 32 GB | 16 GB |
| Usable to GPU◐ | 32 GB | 16 GB |
| Bandwidth○ | 1792 GB/s | 672 GB/s |
| Memory type○ | GDDR7 | GDDR6X |
| Bus width○ | 512-bit | 256-bit |
| Tensor FP16 (dense)○ | 419 TFLOPS | 176 TFLOPS |
| TDP○ | 575 W | 285 W |
| MSRP at launch○ | $1999 | $799 |
| Models that fit◐ | 1954 | 1688 |
Same model, both cards
| Model | Quant | Params○ | GeForce RTX 5090◐ | GeForce RTX 4070 Ti Super◐ | Difference◐ |
|---|---|---|---|---|---|
| Qwen3-Coder-30B-A3B-InstructMoE | Q6_K | 30.5B | 119 | 63 | +88% |
| Qwen3.6-27B | Q6_K | 27.8B | 54 | 35 | +54% |
| Qwen3.6-35B-A3BMoE | UD-Q6_K | 36.0B | 204 | 140 | +45% |
| Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP | IQ3_M | 27.8B | 44 | 35 | +28% |
| Qwen3.8-27B | Q6_K | 27.8B | 54 | 35 | +54% |
| Qwen3.5-9B | BF16 | 9.7B | 69 | 47 | +46% |
| gemma-4-26B-A4B-itMoE | Q8_0 | 26.5B | 46 | 35 | +34% |
| gemma-4-12B-it | BF16 | 12.0B | 51 | 39 | +31% |
| nemotron-3.5-asr-streaming-0.6b | F32 | 638M | 393 | 180 | +118% |
| Qwen3.5-4B | BF16 | 4.7B | 131 | 52 | +150% |
| gemma-4-12B-it-qat-q4_0-unquantized | Q4_0 | 12.0B | 132 | 53 | +150% |
| Qwythos-9B-Claude-Mythos-5-1M | Q8_0 | 9.4B | 66 | 36 | +82% |
Token rates are modeled from memory bandwidth, so on models both cards can hold the ratio tracks bandwidth closely. That is the honest shape of the answer: for inference, capacity decides what you can run and bandwidth decides how fast. Neither is teraflops.
How to read this
We earn nothing from either of these cards. If both hold the models you care about, the faster one wins on bandwidth alone. If one holds a model the other cannot, that difference usually matters far more than any percentage of token rate — a model that does not fit runs 5–20× slower once it starts spilling to system memory, not slightly slower.