Sao10K · text

MN-12B-Lyra-v4

Sao10K/MN-12B-Lyra-v4

MN-12B-Lyra-v4 at Q4_K_M is exactly 7,477,207,776 bytes (6.96 GiB / 7.48 GB) — an effective 4.884 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
12.2B
Architecture
llama
Context
native (config.json)
License
cc-by-nc-4.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M4.13 GiB4,435,026,6562.897bartowski
Q2_K4.46 GiB4,791,050,9763.129bartowski
Q2_K4.46 GiB4,791,051,1363.129mradermacher
IQ3_XS4.94 GiB5,306,491,6163.466bartowski
Q2_K_L5.07 GiB5,446,410,9763.558bartowski
Q3_K_S5.15 GiB5,534,229,2163.615bartowski
Q3_K_S5.15 GiB5,534,229,3763.615mradermacher
IQ3_M5.33 GiB5,722,235,6163.738bartowski
Q3_K_M5.67 GiB6,083,093,2163.973bartowski
Q3_K_M5.67 GiB6,083,093,3763.973mradermacher
Q3_K_L6.11 GiB6,561,506,0164.286bartowski
Q3_K_L6.11 GiB6,561,506,1764.286mradermacher
IQ4_XS6.28 GiB6,742,713,0564.404bartowski
IQ4_XS6.33 GiB6,800,057,2164.442mradermacher
Q4_06.61 GiB7,094,641,3764.634bartowski
Q4_K_S6.63 GiB7,120,200,4164.651bartowski
Q4_K_S6.63 GiB7,120,200,5764.651mradermacher
Q4_K_M6.96 GiB7,477,207,7764.884bartowski
Q4_K_M6.96 GiB7,477,207,9364.884mradermacher
Q4_K_L7.43 GiB7,975,281,3765.209bartowski
Q5_K_S7.93 GiB8,518,738,6565.564bartowski
Q5_K_S7.93 GiB8,518,738,8165.564mradermacher
Q5_K_M8.13 GiB8,727,634,6565.701bartowski
Q5_K_M8.13 GiB8,727,634,8165.701mradermacher
Q5_K_L8.51 GiB9,141,822,1765.971bartowski
Q6_K9.37 GiB10,056,213,2166.569bartowski
Q6_K9.37 GiB10,056,213,3766.569mradermacher
Q6_K_L9.67 GiB10,381,271,7766.781bartowski
Q8_012.13 GiB13,022,372,5768.506bartowski
Q8_012.13 GiB13,022,372,7368.506mradermacher
F1622.82 GiB24,504,279,48816.006Lewdiculous
BF1622.82 GiB24,504,279,48816.006Lewdiculous
F1622.82 GiB24,504,279,48816.006bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 6.42 GiB. The real file is 6.96 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does MN-12B-Lyra-v4 need?
Q4_K_M is exactly 7,477,207,776 bytes (6.96 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of MN-12B-Lyra-v4 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.