Chat-Error · text

Kimiko-v2-13B

Chat-Error/Kimiko-v2-13B

Kimiko-v2-13B at Q4_K_M is exactly 7,865,956,288 bytes (7.33 GiB / 7.87 GB) — an effective 4.835 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
13.0B
Architecture
llama
Context
native (config.json)
License
creativeml-openrail-m

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K5.06 GiB5,429,348,2883.337TheBloke
Q3_K_S5.27 GiB5,658,980,2883.478TheBloke
Q3_K_M5.90 GiB6,337,769,4083.895TheBloke
Q3_K_L6.45 GiB6,929,559,4884.259TheBloke
Q4_06.86 GiB7,365,834,6884.527TheBloke
Q4_K_S6.91 GiB7,414,331,3284.557TheBloke
Q4_K_M7.33 GiB7,865,956,2884.835TheBloke
Q5_K_S8.36 GiB8,972,285,8885.515TheBloke
Q5_08.36 GiB8,972,285,8885.515TheBloke
Q5_K_M8.60 GiB9,229,924,2885.673TheBloke
Q6_K9.95 GiB10,679,140,2886.564TheBloke
Q8_012.88 GiB13,831,319,4888.501TheBloke

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 6.82 GiB. The real file is 7.33 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Kimiko-v2-13B need?
Q4_K_M is exactly 7,865,956,288 bytes (7.33 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Kimiko-v2-13B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.