baidu · text

ERNIE-4.5-21B-A3B-Thinking

baidu/ERNIE-4.5-21B-A3B-Thinking

ERNIE-4.5-21B-A3B-Thinking at Q4_K_M is exactly 13,331,018,624 bytes (12.42 GiB / 13.33 GB) — an effective 4.886 bits per weight, not the nominal 4. Its KV cache at 32K is 1.75 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
21.8B
Architecture
ernie4_5-moe
28 layers
Context
131,072
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S4.17 GiB4,476,480,0321.641mradermacher
I1-IQ1_M4.63 GiB4,969,720,3521.822mradermacher
IQ2_XXS5.19 GiB5,570,050,3682.042bartowski
I1-IQ2_XXS5.39 GiB5,791,787,5522.123mradermacher
IQ2_XS5.91 GiB6,346,979,6482.326bartowski
IQ2_S5.93 GiB6,369,917,2482.335bartowski
I1-IQ2_XS6.01 GiB6,455,175,7122.366mradermacher
UD-TQ1_06.05 GiB6,497,323,9042.382unsloth
I1-IQ2_S6.06 GiB6,510,533,1522.386mradermacher
UD-IQ1_S6.53 GiB7,010,470,7842.570unsloth
IQ2_M6.67 GiB7,164,541,2482.626bartowski
I1-IQ2_M6.68 GiB7,168,186,9122.627mradermacher
UD-IQ1_M6.75 GiB7,243,082,6242.655unsloth
I1-Q2_K_S6.94 GiB7,448,599,0722.730mradermacher
UD-IQ2_XXS7.23 GiB7,767,206,7842.847unsloth
UD-IQ2_M7.48 GiB8,033,405,8242.945unsloth
I1-Q2_K7.50 GiB8,053,066,2722.952mradermacher
Q2_K7.54 GiB8,091,875,6482.966bartowski
Q2_K_L7.59 GiB8,152,599,4242.988unsloth
Q2_K7.59 GiB8,152,599,4242.988unsloth
Q2_K_L7.60 GiB8,155,998,5282.990bartowski
I1-IQ3_XXS7.88 GiB8,456,092,1923.099mradermacher
I1-IQ3_XS8.37 GiB8,983,882,2723.293mradermacher
IQ3_XXS8.38 GiB8,993,098,0483.296bartowski
IQ3_XS8.71 GiB9,350,863,1683.428bartowski
I1-Q3_K_S8.85 GiB9,500,264,9923.482mradermacher
I1-IQ3_S8.85 GiB9,505,139,2323.484mradermacher
UD-IQ3_XXS8.89 GiB9,548,209,0243.500unsloth
I1-IQ3_M8.94 GiB9,602,624,0323.520mradermacher
Q3_K_S8.98 GiB9,639,611,2643.533unsloth
Q3_K_S9.17 GiB9,850,042,6883.611bartowski
IQ3_M9.59 GiB10,293,598,5283.773bartowski
Q3_K_M9.59 GiB10,297,858,3683.775bartowski
I1-Q3_K_M9.75 GiB10,468,579,8723.837mradermacher
Q3_K_M9.80 GiB10,524,818,3043.858unsloth
Q3_K_L9.93 GiB10,663,754,0483.909bartowski
I1-Q3_K_L10.59 GiB11,371,665,9524.168mradermacher
I1-IQ4_XS10.89 GiB11,695,290,9124.287mradermacher
IQ4_XS11.01 GiB11,823,025,0244.334unsloth
IQ4_XS11.14 GiB11,958,008,1284.383bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.22 GiB0.22 GiB28 / 0 / 0
8,1920.44 GiB0.44 GiB28 / 0 / 0
16,3840.88 GiB0.88 GiB28 / 0 / 0
32,7681.75 GiB1.75 GiB28 / 0 / 0
65,5363.50 GiB3.50 GiB28 / 0 / 0
131,0727.00 GiB7.00 GiB28 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 11.43 GiB. The real file is 12.42 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
28
Attention heads
20
KV heads
4
Head dim
128
Hidden size
2560
Vocab
103,424
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does ERNIE-4.5-21B-A3B-Thinking need?
Q4_K_M is exactly 13,331,018,624 bytes (12.42 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is ERNIE-4.5-21B-A3B-Thinking's KV cache?
1.75 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of ERNIE-4.5-21B-A3B-Thinking should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.