solar-pro-preview-instruct
upstage/solar-pro-preview-instructsolar-pro-preview-instruct at Q4_K_M is exactly 13,310,219,072 bytes (12.40 GiB / 13.31 GB) — an effective 4.809 bits per weight, not the nominal 4.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| IQ1_S | 4.46 GiB | 4,786,663,232 | 1.730 | — | MaziyarPanahi |
| IQ1_M | 4.87 GiB | 5,231,488,832 | 1.890 | — | MaziyarPanahi |
| IQ2_XS | 6.16 GiB | 6,618,394,432 | 2.392 | — | MaziyarPanahi |
| Q2_K | 7.65 GiB | 8,213,924,672 | 2.968 | — | MaziyarPanahi |
| IQ3_XS | 8.50 GiB | 9,125,197,632 | 3.297 | — | MaziyarPanahi |
| Q3_K_S | 8.92 GiB | 9,580,672,832 | 3.462 | — | MaziyarPanahi |
| Q3_K_M | 9.95 GiB | 10,686,592,832 | 3.861 | 579 | MaziyarPanahi |
| Q3_K_L | 10.84 GiB | 11,635,226,432 | 4.204 | — | MaziyarPanahi |
| IQ4_XS | 11.06 GiB | 11,878,032,192 | 4.292 | 579 | MaziyarPanahi |
| Q4_K_S | 11.73 GiB | 12,594,238,272 | 4.551 | — | MaziyarPanahi |
| Q4_K_M | 12.40 GiB | 13,310,219,072 | 4.809 | 579 | MaziyarPanahi |
| Q5_K_S | 14.20 GiB | 15,246,070,592 | 5.509 | — | MaziyarPanahi |
| Q5_K_M | 14.59 GiB | 15,663,862,592 | 5.660 | 579 | MaziyarPanahi |
| Q6_K | 16.92 GiB | 18,164,608,832 | 6.564 | 579 | MaziyarPanahi |
| Q8_0 | 21.91 GiB | 23,526,487,872 | 8.501 | 579 | MaziyarPanahi |
KV cache by context
This model declares a 2,047-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.
A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 11.60 GiB. The real file is 12.40 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Questions people ask
- How much VRAM does solar-pro-preview-instruct need?
- Q4_K_M is exactly 13,310,219,072 bytes (12.40 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Which quantization of solar-pro-preview-instruct should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.