SpydazWeb_AI_CyberTron_Ultra_7b
LeroyDyer/SpydazWeb_AI_CyberTron_Ultra_7bSpydazWeb_AI_CyberTron_Ultra_7b at Q4_K_M is exactly 4,368,439,488 bytes (4.07 GiB / 4.37 GB) — an effective 4.826 bits per weight, not the nominal 4.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| Q2_K | 2.53 GiB | 2,719,242,432 | 3.004 | — | mradermacher |
| IQ3_XS | 2.81 GiB | 3,018,815,680 | 3.335 | — | mradermacher |
| Q3_K_S | 2.95 GiB | 3,164,567,744 | 3.496 | — | mradermacher |
| IQ3_S | 2.96 GiB | 3,182,393,536 | 3.516 | — | mradermacher |
| IQ3_M | 3.06 GiB | 3,284,891,840 | 3.629 | — | mradermacher |
| Q3_K_M | 3.28 GiB | 3,518,986,432 | 3.888 | — | mradermacher |
| Q3_K_L | 3.56 GiB | 3,822,024,896 | 4.222 | — | mradermacher |
| IQ4_XS | 3.67 GiB | 3,944,388,800 | 4.357 | — | mradermacher |
| Q4_K_S | 3.86 GiB | 4,140,374,208 | 4.574 | — | mradermacher |
| Q4_K_M | 4.07 GiB | 4,368,439,488 | 4.826 | — | mradermacher |
| Q5_K_S | 4.65 GiB | 4,997,716,160 | 5.521 | — | mradermacher |
| Q5_K_M | 4.78 GiB | 5,131,409,600 | 5.669 | — | mradermacher |
| Q6_K | 5.53 GiB | 5,942,065,344 | 6.564 | — | mradermacher |
| Q8_0 | 7.17 GiB | 7,695,857,856 | 8.502 | — | mradermacher |
KV cache by context
This model declares a 4,096-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.
A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 3.79 GiB. The real file is 4.07 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Questions people ask
- How much VRAM does SpydazWeb_AI_CyberTron_Ultra_7b need?
- Q4_K_M is exactly 4,368,439,488 bytes (4.07 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Which quantization of SpydazWeb_AI_CyberTron_Ultra_7b should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.