TheDrummer · text

Behemoth-X-123B-v2

TheDrummer/Behemoth-X-123B-v2

Behemoth-X-123B-v2 at Q4_K_M is exactly 73,219,623,488 bytes (68.19 GiB / 73.22 GB) — an effective 4.777 bits per weight, not the nominal 4. Its KV cache at 32K is 11.00 GiB.

From the file· summed from 2 file(s)From the file· KV per layer
Parameters
123B
Architecture
llama
88 layers
Context
131,072
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S24.18 GiB25,959,778,6881.694bartowski
IQ1_M26.44 GiB28,386,314,6241.852bartowski
IQ2_XXS30.20 GiB32,430,541,1842.116bartowski
IQ2_XS33.60 GiB36,081,158,5282.354bartowski
IQ2_S35.75 GiB38,384,224,6402.505bartowski
IQ2_M38.76 GiB41,619,605,8882.716bartowski
Q2_K42.09 GiB45,196,298,6242.949bartowski
Q2_K_L42.46 GiB45,589,514,6242.975bartowski
IQ3_XXS43.78 GiB47,009,024,3843.067bartowski
IQ3_XS2 shards46.70 GiB50,142,169,6643.272bartowski
Q3_K_S2 shards49.22 GiB52,849,855,0403.448bartowski
IQ3_M2 shards51.48 GiB55,276,390,9443.607bartowski
Q3_K_M2 shards55.04 GiB59,102,775,8723.856bartowski
Q3_K_L2 shards60.12 GiB64,554,322,4964.212bartowski
IQ4_XS2 shards60.94 GiB65,434,339,9044.269bartowski
IQ4_NL2 shards64.46 GiB69,218,650,6884.516bartowski
Q4_02 shards64.56 GiB69,322,459,7124.523bartowski
Q4_K_S2 shards64.79 GiB69,570,972,2244.539bartowski
Q4_K_M2 shards68.19 GiB73,219,623,4884.777bartowski
Q4_K_L2 shards68.47 GiB73,518,467,6484.797bartowski
Q4_12 shards71.45 GiB76,718,066,2085.006bartowski
Q5_K_S3 shards78.56 GiB84,355,893,9205.504bartowski
Q5_K_M3 shards80.55 GiB86,488,304,2885.643bartowski
Q6_K3 shards93.68 GiB100,586,277,5686.563bartowski
Q8_04 shards121.33 GiB130,280,376,8648.501bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.38 GiB1.38 GiB88 / 0 / 0
8,1922.75 GiB2.75 GiB88 / 0 / 0
16,3845.50 GiB5.50 GiB88 / 0 / 0
32,76811.00 GiB11.00 GiB88 / 0 / 0
65,53622.00 GiB22.00 GiB88 / 0 / 0
131,07244.00 GiB44.00 GiB88 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 64.23 GiB. The real file is 68.19 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
88
Attention heads
96
KV heads
8
Head dim
128
Hidden size
12288
Vocab
32,768
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Behemoth-X-123B-v2 need?
Q4_K_M is exactly 73,219,623,488 bytes (68.19 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Behemoth-X-123B-v2's KV cache?
11.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Behemoth-X-123B-v2 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.