TheDrummer · text

Snowpiercer-15B-v4

TheDrummer/Snowpiercer-15B-v4

Snowpiercer-15B-v4 at Q4_K_M is exactly 9,112,538,240 bytes (8.49 GiB / 9.11 GB) — an effective 4.868 bits per weight, not the nominal 4. Its KV cache at 32K is 6.25 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
15.0B
Architecture
llama
50 layers
Context
65,536
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_S4.64 GiB4,981,107,8402.661bartowski
IQ2_M4.98 GiB5,352,369,2802.860bartowski
Q2_K5.40 GiB5,794,163,8403.095bartowski
IQ3_XXS5.58 GiB5,992,328,3203.201bartowski
IQ3_XS5.98 GiB6,424,865,9203.433bartowski
Q2_K_L6.01 GiB6,449,523,8403.446bartowski
Q3_K_S6.25 GiB6,706,097,2803.583bartowski
IQ3_M6.46 GiB6,938,668,1603.707bartowski
Q3_K_M6.89 GiB7,396,437,1203.952bartowski
Q3_K_L7.44 GiB7,990,193,2804.269bartowski
IQ4_XS7.64 GiB8,199,662,7204.381bartowski
Q4_08.04 GiB8,633,183,3604.612bartowski
IQ4_NL8.05 GiB8,638,426,2404.615bartowski
Q4_K_S8.07 GiB8,663,329,9204.628bartowski
Q4_K_M8.49 GiB9,112,538,2404.868bartowski
Q4_18.85 GiB9,499,569,2805.075bartowski
Q4_K_L8.95 GiB9,610,611,8405.135bartowski
Q5_K_S9.68 GiB10,393,480,3205.553bartowski
Q5_K_M9.92 GiB10,654,600,3205.692bartowski
Q5_K_L10.31 GiB11,068,787,8405.913bartowski
Q6_K11.45 GiB12,293,041,2806.568bartowski
Q6_K_L11.75 GiB12,618,099,8406.741bartowski
Q8_014.83 GiB15,919,475,8408.505bartowski
BF1627.90 GiB29,957,286,75216.005bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.78 GiB0.78 GiB50 / 0 / 0
8,1921.56 GiB1.56 GiB50 / 0 / 0
16,3843.13 GiB3.13 GiB50 / 0 / 0
32,7686.25 GiB6.25 GiB50 / 0 / 0
65,53612.50 GiB12.50 GiB50 / 0 / 0
131,07225.00 GiB25.00 GiB50 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 7.84 GiB. The real file is 8.49 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
50
Attention heads
32
KV heads
8
Head dim
128
Hidden size
5120
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Snowpiercer-15B-v4 need?
Q4_K_M is exactly 9,112,538,240 bytes (8.49 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Snowpiercer-15B-v4's KV cache?
6.25 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Snowpiercer-15B-v4 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.