Triangle104 · text

Mistral_Sunair-V1.0

Triangle104/Mistral_Sunair-V1.0

Mistral_Sunair-V1.0 at Q4_K_M is exactly 7,477,209,152 bytes (6.96 GiB / 7.48 GB) — an effective 4.884 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
12.2B
Architecture
llama
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S2.79 GiB2,999,216,4161.959mradermacher
I1-IQ1_M3.00 GiB3,221,629,2162.104mradermacher
I1-IQ2_XXS3.35 GiB3,592,317,2162.346mradermacher
I1-IQ2_XS3.65 GiB3,915,082,0162.557mradermacher
I1-IQ2_S3.85 GiB4,138,477,8562.703mradermacher
I1-IQ2_M4.13 GiB4,435,028,2562.897mradermacher
Q2_K4.46 GiB4,791,052,3523.129mradermacher
I1-Q2_K4.46 GiB4,791,052,5763.129mradermacher
I1-IQ3_XXS4.61 GiB4,945,389,8563.230mradermacher
I1-IQ3_XS4.94 GiB5,306,493,2163.466mradermacher
Q3_K_S5.15 GiB5,534,230,5923.615mradermacher
I1-Q3_K_S5.15 GiB5,534,230,8163.615mradermacher
I1-IQ3_S5.18 GiB5,562,083,6163.633mradermacher
I1-IQ3_M5.33 GiB5,722,237,2163.738mradermacher
Q3_K_M5.67 GiB6,083,094,5923.973mradermacher
I1-Q3_K_M5.67 GiB6,083,094,8163.973mradermacher
Q3_K_L6.11 GiB6,561,507,3924.286mradermacher
I1-Q3_K_L6.11 GiB6,561,507,6164.286mradermacher
I1-IQ4_XS6.28 GiB6,742,714,6564.404mradermacher
IQ4_XS6.33 GiB6,800,058,4324.442mradermacher
I1-Q4_06.61 GiB7,094,642,9764.634mradermacher
Q4_K_S6.63 GiB7,120,201,7924.651mradermacher
I1-Q4_K_S6.63 GiB7,120,202,0164.651mradermacher
Q4_K_M6.96 GiB7,477,209,1524.884mradermacher
I1-Q4_K_M6.96 GiB7,477,209,3764.884mradermacher
Q5_K_S7.93 GiB8,518,740,0325.564mradermacher
I1-Q5_K_S7.93 GiB8,518,740,2565.564mradermacher
Q5_K_M8.13 GiB8,727,636,0325.701mradermacher
I1-Q5_K_M8.13 GiB8,727,636,2565.701mradermacher
Q6_K9.37 GiB10,056,214,5926.569mradermacher
I1-Q6_K9.37 GiB10,056,214,8166.569mradermacher
Q8_012.13 GiB13,022,373,9528.506mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 6.42 GiB. The real file is 6.96 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Mistral_Sunair-V1.0 need?
Q4_K_M is exactly 7,477,209,152 bytes (6.96 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Mistral_Sunair-V1.0 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.