Best local vision-language models
Models that accept images alongside text. Remember these need their vision projector loaded too, and that images consume context tokens quickly.
From the file· live filter over real dataFrom the file· 40 models
How this is ranked
Ranked by downloads among vision-language models with published quantizations.
Best local vision-language models
Best local models for codingBest local reasoning modelsBest models for an 8GB cardBest models for long contextBest permissively licensed modelsBest models under 4GBLargest models with a runnable quantizationBest local speech recognition modelsBest local text-to-speech modelsBest local image generation models