Best local AI models for 6GB VRAM

Ranked by what actually fits at 32K context, computed from real file bytes.

Nothing in the current index fits 6GB at 32K context with f16 KV.

From the file· fit from summed bytesFrom the file· KV per layer

Fits in 6GB at 32K context

largest quantization that fits, per model
ModelModalityBest quantParamsTotalHeadroom
Spec sheetPredictedwhat these mean

This page models a generic 6GB accelerator, so it answers what fits rather than how fast it runs. For tokens per second you need a specific card — pick one from hardware, where bandwidth is known.