Best local AI models for 128GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

A 128GB card gives you about 119.04 GiB to work with after driver overhead. 2 indexed models fit at 32K context — the largest being HunyuanImage-2.1 at 17.5B parameters in IQ4_NL.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 128GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
HunyuanImage-2.1image generationIQ4_NL17.5B79.47 GiB39.57 GiB
Janus-Pro-7Bimage generationF167.4B28.70 GiB90.34 GiB
Spec sheetPredictedwhat these mean

This page models a generic 128GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.