Quantization publisher

NeuralNet-Hub

NeuralNet-Hub publishes 14 quantizations across 1 models in our index, averaging 5.175 effective bits per weight. Their files differ in size from other publishers' builds of the same nominal quantization on 4 of the pairs we can compare — the same label does not mean the same file.

From the file· summed file bytes
Repositories
1
Quantizations
14
Models covered
1
Avg effective bpw
5.175
across their files

Same model, same quant label, different bytes

largest disagreements first
ModelQuantNeuralNet-HubvsTheirsDifference
openchat-3.6-8b-20240522Q2_K2.96 GiBlegraphista2.96 GiB-0.0%
openchat-3.6-8b-20240522Q3_K_L4.03 GiBlegraphista4.03 GiB-0.0%
openchat-3.6-8b-20240522Q3_K_S3.41 GiBlegraphista3.41 GiB-0.0%
openchat-3.6-8b-20240522Q4_K_S4.37 GiBlegraphista4.37 GiB-0.0%

A quantization label describes a target, not a recipe. Publishers make different choices about which tensors to keep at higher precision, and some apply an importance matrix while others don't — so two files both honestly labelled the same thing can differ measurably in size and in quality.

Models they publish

ModelQuantizationsSmallest
openchat-3.6-8b-20240522142.96 GiB