speakleash · text

Bielik-11B-v2.2-Instruct

speakleash/Bielik-11B-v2.2-Instruct

Bielik-11B-v2.2-Instruct at Q4_K_M is exactly 6,724,050,432 bytes (6.26 GiB / 6.72 GB) — an effective 4.816 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
11.2B
Architecture
llama
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M3.56 GiB3,823,567,1362.739bartowski
Q2_K3.88 GiB4,164,337,9522.983bartowski
Q2_K_L4.00 GiB4,292,849,9523.075bartowski
IQ3_XS4.31 GiB4,627,738,9123.315bartowski
Q3_K_S4.52 GiB4,852,724,0003.476bartowski
IQ3_M4.69 GiB5,038,780,7043.609bartowski
Q3_K_M5.03 GiB5,404,995,8723.872bartowski
Q3_K_L5.48 GiB5,880,000,8004.212bartowski
IQ4_XS5.59 GiB6,006,415,6484.302bartowski
Q4_05.91 GiB6,340,567,3284.542bartowski
Q4_K_S5.93 GiB6,364,684,5764.559bartowski
Q4_K_M6.26 GiB6,724,050,4324.816speakleash
Q4_K_M6.26 GiB6,724,051,2324.816bartowski
Q4_K_L6.35 GiB6,821,720,3524.886bartowski
Q5_K_S7.17 GiB7,698,145,5685.514bartowski
Q5_K_M7.36 GiB7,907,040,7685.664speakleash
Q5_K_M7.36 GiB7,907,041,5685.664bartowski
Q5_K_L7.44 GiB7,988,261,1525.722bartowski
Q6_K8.53 GiB9,163,968,0006.564speakleash
Q6_K8.53 GiB9,163,968,8006.564bartowski
Q6_K_L8.59 GiB9,227,710,7526.610bartowski
Q8_011.05 GiB11,868,810,7528.501speakleash
Q8_011.05 GiB11,868,811,5528.501bartowski
F1620.80 GiB22,339,170,30416.001bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 5.85 GiB. The real file is 6.26 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Bielik-11B-v2.2-Instruct need?
Q4_K_M is exactly 6,724,050,432 bytes (6.26 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Bielik-11B-v2.2-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.