Quantization publisher

LiteLLMs

LiteLLMs publishes 14 quantizations across 1 models in our index, averaging 5.095 effective bits per weight. Their files differ in size from other publishers' builds of the same nominal quantization on 5 of the pairs we can compare — the same label does not mean the same file.

From the file· summed file bytes
Repositories
1
Quantizations
14
Models covered
1
Avg effective bpw
5.095
across their files

Same model, same quant label, different bytes

largest disagreements first
ModelQuantLiteLLMsvsTheirsDifference
Meta-Llama-3-70B-InstructQ4_K_M39.61 GiBlmstudio-community39.60 GiB0.0%
Meta-Llama-3-70B-InstructQ4_K_M39.61 GiBqwp4w3hyb39.60 GiB0.0%
Meta-Llama-3-70B-InstructQ4_K_S37.58 GiBqwp4w3hyb37.58 GiB0.0%
Meta-Llama-3-70B-InstructQ5_K_M46.53 GiBqwp4w3hyb46.52 GiB0.0%
Meta-Llama-3-70B-InstructQ5_K_S45.32 GiBqwp4w3hyb45.32 GiB0.0%

A quantization label describes a target, not a recipe. Publishers make different choices about which tensors to keep at higher precision, and some apply an importance matrix while others don't — so two files both honestly labelled the same thing can differ measurably in size and in quality.

Models they publish

ModelQuantizationsSmallest
Meta-Llama-3-70B-Instruct1424.56 GiB