SciPhi-Self-RAG-Mistral-7B-32k
SciPhi/SciPhi-Self-RAG-Mistral-7B-32kSciPhi-Self-RAG-Mistral-7B-32k at Q4_K_M is exactly 4,368,530,080 bytes (4.07 GiB / 4.37 GB) — an effective 4.826 bits per weight, not the nominal 4.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| I1-IQ1_S | 1.50 GiB | 1,612,169,408 | 1.781 | — | mradermacher |
| I1-IQ1_M | 1.63 GiB | 1,754,513,600 | 1.938 | — | mradermacher |
| I1-IQ2_XXS | 1.85 GiB | 1,991,753,920 | 2.200 | — | mradermacher |
| I1-IQ2_XS | 2.05 GiB | 2,198,323,392 | 2.429 | — | mradermacher |
| I1-IQ2_S | 2.15 GiB | 2,310,994,624 | 2.553 | — | mradermacher |
| I1-IQ2_M | 2.33 GiB | 2,500,786,880 | 2.763 | — | mradermacher |
| I1-Q2_K | 2.53 GiB | 2,719,318,720 | 3.004 | — | mradermacher |
| I1-IQ3_XXS | 2.63 GiB | 2,827,418,304 | 3.123 | — | mradermacher |
| I1-IQ3_XS | 2.81 GiB | 3,018,898,624 | 3.335 | — | mradermacher |
| Q2_K | 2.87 GiB | 3,083,173,536 | 3.406 | — | TheBloke |
| Q3_K_S | 2.95 GiB | 3,164,649,632 | 3.496 | — | TheBloke |
| I1-Q3_K_S | 2.95 GiB | 3,164,650,688 | 3.496 | — | mradermacher |
| I1-IQ3_S | 2.96 GiB | 3,182,476,480 | 3.516 | — | mradermacher |
| I1-IQ3_M | 3.06 GiB | 3,284,974,784 | 3.629 | — | mradermacher |
| Q3_K_M | 3.28 GiB | 3,519,068,320 | 3.888 | — | TheBloke |
| I1-Q3_K_M | 3.28 GiB | 3,519,069,376 | 3.888 | — | mradermacher |
| Q3_K_L | 3.56 GiB | 3,822,106,784 | 4.222 | — | TheBloke |
| I1-Q3_K_L | 3.56 GiB | 3,822,107,840 | 4.222 | — | mradermacher |
| I1-IQ4_XS | 3.64 GiB | 3,907,778,240 | 4.317 | — | mradermacher |
| Q4_0 | 3.83 GiB | 4,109,007,520 | 4.539 | — | TheBloke |
| I1-Q4_0 | 3.84 GiB | 4,123,688,640 | 4.555 | — | mradermacher |
| Q4_K_S | 3.86 GiB | 4,140,464,800 | 4.574 | — | TheBloke |
| I1-Q4_K_S | 3.86 GiB | 4,140,465,856 | 4.574 | — | mradermacher |
| Q4_K_M | 4.07 GiB | 4,368,530,080 | 4.826 | — | TheBloke |
| I1-Q4_K_M | 4.07 GiB | 4,368,531,136 | 4.826 | — | mradermacher |
| Q5_0 | 4.65 GiB | 4,997,814,944 | 5.521 | — | TheBloke |
| Q5_K_S | 4.65 GiB | 4,997,814,944 | 5.521 | — | TheBloke |
| I1-Q5_K_S | 4.65 GiB | 4,997,816,000 | 5.521 | — | mradermacher |
| Q5_K_M | 4.78 GiB | 5,131,508,384 | 5.669 | — | TheBloke |
| I1-Q5_K_M | 4.78 GiB | 5,131,509,440 | 5.669 | — | mradermacher |
| Q6_K | 5.53 GiB | 5,942,172,832 | 6.564 | — | TheBloke |
| I1-Q6_K | 5.53 GiB | 5,942,173,888 | 6.564 | — | mradermacher |
| Q8_0 | 7.17 GiB | 7,695,997,088 | 8.502 | — | TheBloke |
KV cache by context
This model declares a 4,096-token sliding window, but we could not establish which layers use it. Its architecture publishes the layout as a per-layer array inside the model file rather than as a period in config.json, and we have not yet ingested that array.
A flat context × layers × heads figure would be substantially too high, so we are not showing one. This is tracked as a known gap rather than filled with a guess.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 3.79 GiB. The real file is 4.07 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.
Architecture
Questions people ask
- How much VRAM does SciPhi-Self-RAG-Mistral-7B-32k need?
- Q4_K_M is exactly 4,368,530,080 bytes (4.07 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- Which quantization of SciPhi-Self-RAG-Mistral-7B-32k should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.