dphn · text · mixture of experts

dolphin-2.9.1-mixtral-1x22b

dphn/dolphin-2.9.1-mixtral-1x22b

dolphin-2.9.1-mixtral-1x22b at IQ1_S is exactly 4,826,066,272 bytes (4.49 GiB / 4.83 GB) — an effective 1.736 bits per weight, not the nominal 1. Its KV cache at 32K is 7.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
22.2B
total, not active
Architecture
llama
56 layers
Context
65,536
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S4.49 GiB4,826,066,2721.736legraphista
I1-IQ1_S4.49 GiB4,826,066,7201.736mradermacher
IQ1_M4.90 GiB5,263,715,6801.894legraphista
I1-IQ1_M4.90 GiB5,263,716,1281.894mradermacher
IQ2_XXS5.58 GiB5,993,131,3602.156legraphista
I1-IQ2_XXS5.58 GiB5,993,131,8082.156mradermacher
IQ2_XS6.19 GiB6,642,724,1922.390legraphista
I1-IQ2_XS6.19 GiB6,642,724,6402.390mradermacher
IQ2_S6.55 GiB7,031,530,0482.530legraphista
I1-IQ2_S6.55 GiB7,031,530,4962.530mradermacher
IQ2_M7.09 GiB7,615,062,5922.740legraphista
I1-IQ2_M7.09 GiB7,615,063,0402.740mradermacher
Q2_K_S7.12 GiB7,645,979,5842.751legraphista
Q2_K7.70 GiB8,268,047,2962.974legraphista
I1-Q2_K7.70 GiB8,268,047,7442.974mradermacher
IQ3_XXS8.00 GiB8,594,956,8643.092legraphista
I1-IQ3_XXS8.00 GiB8,594,957,3123.092mradermacher
IQ3_XS8.54 GiB9,171,572,8963.299legraphista
I1-IQ3_XS8.54 GiB9,171,573,3443.299mradermacher
Q3_K_S8.97 GiB9,636,747,4243.467legraphista
I1-Q3_K_S8.97 GiB9,636,747,8723.467mradermacher
IQ3_S9.02 GiB9,683,540,1283.484legraphista
I1-IQ3_S9.02 GiB9,683,540,5763.484mradermacher
IQ3_M9.37 GiB10,057,881,7603.618legraphista
I1-IQ3_M9.37 GiB10,057,882,2083.618mradermacher
Q3_K10.01 GiB10,752,301,2163.868legraphista
I1-Q3_K_M10.01 GiB10,752,301,6643.868mradermacher
Q3_K_L10.92 GiB11,725,904,0324.218legraphista
I1-Q3_K_L10.92 GiB11,725,904,4804.218mradermacher
IQ4_XS11.11 GiB11,930,291,5844.292legraphista
I1-IQ4_XS11.11 GiB11,930,292,0324.292mradermacher
IQ4_NL11.74 GiB12,608,048,8964.536legraphista
I1-Q4_011.74 GiB12,608,049,3444.536mradermacher
Q4_K_S11.79 GiB12,655,234,8164.553legraphista
I1-Q4_K_S11.79 GiB12,655,235,2644.553mradermacher
Q4_K12.42 GiB13,336,088,3204.798legraphista
I1-Q4_K_M12.42 GiB13,336,088,7684.798mradermacher
Q5_K_S14.27 GiB15,319,077,8565.511legraphista
I1-Q5_K_S14.27 GiB15,319,078,5925.511mradermacher
Q5_K14.64 GiB15,716,815,8405.654legraphista

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.88 GiB0.88 GiB56 / 0 / 0
8,1921.75 GiB1.75 GiB56 / 0 / 0
16,3843.50 GiB3.50 GiB56 / 0 / 0
32,7687.00 GiB7.00 GiB56 / 0 / 0
65,53614.00 GiB14.00 GiB56 / 0 / 0
131,07228.00 GiB28.00 GiB56 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts IQ1_S at roughly 11.65 GiB. The real file is 4.49 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
56
Attention heads
48
KV heads
8
Head dim
128
Hidden size
6144
Vocab
32,002
Sliding window
none
SWA period
MLA
no
Experts
1
Experts per token
1
use_sliding_window

Questions people ask

How much VRAM does dolphin-2.9.1-mixtral-1x22b need?
IQ1_S is exactly 4,826,066,272 bytes (4.49 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is dolphin-2.9.1-mixtral-1x22b's KV cache?
7.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Is dolphin-2.9.1-mixtral-1x22b a mixture-of-experts model?
Yes — 1 experts, 1 routed per token. Every expert must be resident, but only the routed ones are read per token, which is why its memory requirement and its speed behave very differently.
Which quantization of dolphin-2.9.1-mixtral-1x22b should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.