Quantization publisher
NeuralNet-Hub
NeuralNet-Hub publishes 14 quantizations across 1 models in our index, averaging 5.175 effective bits per weight. Their files differ in size from other publishers' builds of the same nominal quantization on 4 of the pairs we can compare — the same label does not mean the same file.
From the file· summed file bytes
Repositories
1
Quantizations
14
Models covered
1
Avg effective bpw
5.175
across their files
Same model, same quant label, different bytes
largest disagreements first
| Model | Quant | NeuralNet-Hub | vs | Theirs | Difference |
|---|---|---|---|---|---|
| openchat-3.6-8b-20240522 | Q2_K | 2.96 GiB | legraphista | 2.96 GiB | -0.0% |
| openchat-3.6-8b-20240522 | Q3_K_L | 4.03 GiB | legraphista | 4.03 GiB | -0.0% |
| openchat-3.6-8b-20240522 | Q3_K_S | 3.41 GiB | legraphista | 3.41 GiB | -0.0% |
| openchat-3.6-8b-20240522 | Q4_K_S | 4.37 GiB | legraphista | 4.37 GiB | -0.0% |
A quantization label describes a target, not a recipe. Publishers make different choices about which tensors to keep at higher precision, and some apply an importance matrix while others don't — so two files both honestly labelled the same thing can differ measurably in size and in quality.
Models they publish
| Model | Quantizations | Smallest |
|---|---|---|
| openchat-3.6-8b-20240522 | 14 | 2.96 GiB |