sh0ck0r · text

Snowpiercer-15B-v4-heretic

sh0ck0r/Snowpiercer-15B-v4-heretic

Snowpiercer-15B-v4-heretic at Q4_K_M is exactly 9,112,538,592 bytes (8.49 GiB / 9.11 GB) — an effective 4.868 bits per weight, not the nominal 4. Its KV cache at 32K is 6.25 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
15.0B
Architecture
llama
50 layers
Context
65,536
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S3.33 GiB3,574,214,3681.909mradermacher
I1-IQ1_M3.59 GiB3,852,660,4482.058mradermacher
I1-IQ2_XXS4.02 GiB4,316,737,2482.306mradermacher
I1-IQ2_XS4.40 GiB4,720,766,6882.522mradermacher
I1-IQ2_S4.64 GiB4,981,108,4482.661mradermacher
I1-IQ2_M4.98 GiB5,352,369,8882.860mradermacher
I1-Q2_K_S5.05 GiB5,418,151,6482.895mradermacher
Q2_K5.40 GiB5,794,164,1923.095mradermacher
I1-Q2_K5.40 GiB5,794,164,4483.095mradermacher
I1-IQ3_XXS5.58 GiB5,992,328,9283.201mradermacher
I1-IQ3_XS5.98 GiB6,424,866,5283.433mradermacher
Q3_K_S6.25 GiB6,706,097,6323.583mradermacher
I1-Q3_K_S6.25 GiB6,706,097,8883.583mradermacher
I1-IQ3_S6.28 GiB6,740,913,8883.601mradermacher
I1-IQ3_M6.46 GiB6,938,668,7683.707mradermacher
Q3_K_M6.89 GiB7,396,437,4723.952mradermacher
I1-Q3_K_M6.89 GiB7,396,437,7283.952mradermacher
Q3_K_L7.44 GiB7,990,193,6324.269mradermacher
I1-Q3_K_L7.44 GiB7,990,193,8884.269mradermacher
I1-IQ4_XS7.64 GiB8,199,663,3284.381mradermacher
IQ4_XS7.70 GiB8,268,475,8724.418mradermacher
I1-Q4_08.04 GiB8,633,183,9684.612mradermacher
I1-IQ4_NL8.05 GiB8,638,426,8484.615mradermacher
Q4_K_S8.07 GiB8,663,330,2724.628mradermacher
I1-Q4_K_S8.07 GiB8,663,330,5284.628mradermacher
Q4_K_M8.49 GiB9,112,538,5924.868mradermacher
I1-Q4_K_M8.49 GiB9,112,538,8484.868mradermacher
I1-Q4_18.85 GiB9,499,569,8885.075mradermacher
Q5_K_S9.68 GiB10,393,480,6725.553mradermacher
I1-Q5_K_S9.68 GiB10,393,480,9285.553mradermacher
Q5_K_M9.92 GiB10,654,600,6725.692mradermacher
I1-Q5_K_M9.92 GiB10,654,600,9285.692mradermacher
Q6_K11.45 GiB12,293,041,6326.568mradermacher
I1-Q6_K11.45 GiB12,293,041,8886.568mradermacher
Q8_014.83 GiB15,919,476,1928.505mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.78 GiB0.78 GiB50 / 0 / 0
8,1921.56 GiB1.56 GiB50 / 0 / 0
16,3843.13 GiB3.13 GiB50 / 0 / 0
32,7686.25 GiB6.25 GiB50 / 0 / 0
65,53612.50 GiB12.50 GiB50 / 0 / 0
131,07225.00 GiB25.00 GiB50 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 7.84 GiB. The real file is 8.49 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
50
Attention heads
32
KV heads
8
Head dim
128
Hidden size
5120
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Snowpiercer-15B-v4-heretic need?
Q4_K_M is exactly 9,112,538,592 bytes (8.49 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Snowpiercer-15B-v4-heretic's KV cache?
6.25 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Snowpiercer-15B-v4-heretic should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.