LGAI-EXAONE · text

EXAONE-4.0-32B

LGAI-EXAONE/EXAONE-4.0-32B

EXAONE-4.0-32B at Q4_K_M is exactly 19,343,826,720 bytes (18.02 GiB / 19.34 GB) — an effective 4.835 bits per weight, not the nominal 4. Its KV cache at 32K is 2.84 GiB, not the 8.00 GiB a flat formula predicts.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
32.0B
Architecture
exaone4
64 layers
Context
131,072
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XS8.96 GiB9,622,626,5922.405bartowski
IQ2_S9.34 GiB10,025,754,9122.506bartowski
IQ2_M10.15 GiB10,895,089,9522.724bartowski
Q2_K11.11 GiB11,926,462,7522.981bartowski
Q2_K_L11.58 GiB12,438,462,7523.109bartowski
IQ3_XXS11.60 GiB12,455,338,2723.114bartowski
IQ3_XS12.37 GiB13,281,911,0723.320bartowski
Q3_K_S13.00 GiB13,962,830,1123.490bartowski
IQ3_M13.39 GiB14,379,229,4723.594bartowski
Q3_K_M14.43 GiB15,493,751,0723.873708bartowski
Q3_K_L15.64 GiB16,795,951,3924.199bartowski
IQ4_XS16.03 GiB17,212,268,8324.303bartowski
IQ4_XS16.19 GiB17,387,577,1204.346708LGAI-EXAONE
IQ4_NL16.94 GiB18,185,478,4324.546bartowski
Q4_016.96 GiB18,213,658,9124.553708bartowski
Q4_K_S17.03 GiB18,286,403,8724.571bartowski
Q4_K_M18.02 GiB19,343,826,7204.835708LGAI-EXAONE
Q4_K_M18.02 GiB19,343,827,2324.835bartowski
Q4_K_L18.38 GiB19,732,947,2324.933bartowski
Q4_118.73 GiB20,110,926,1125.027bartowski
Q5_K_S20.56 GiB22,078,316,8325.519bartowski
Q5_K_M21.14 GiB22,696,648,4805.674708LGAI-EXAONE
Q5_K_M21.14 GiB22,696,648,9925.674bartowski
Q5_K_L21.44 GiB23,020,232,9925.755bartowski
Q6_K24.46 GiB26,259,021,6006.564708LGAI-EXAONE
Q6_K24.46 GiB26,259,022,1126.564bartowski
Q6_K_L24.69 GiB26,512,974,1126.628bartowski
Q8_031.67 GiB34,009,636,6408.502708LGAI-EXAONE
Q8_031.67 GiB34,009,637,1528.502bartowski
BF162 shards59.62 GiB64,012,017,63216.001LGAI-EXAONE
BF162 shards59.62 GiB64,012,017,88816.001bartowski

KV cache by context

computed per layer — this model uses sliding-window attention
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.00 GiB1.00 GiB16 / 48 / 0
8,1921.34 GiB2.00 GiB1.49×16 / 48 / 0
16,3841.84 GiB4.00 GiB2.17×16 / 48 / 0
32,7682.84 GiB8.00 GiB2.81×16 / 48 / 0
65,5364.84 GiB16.00 GiB3.30×16 / 48 / 0
131,0728.84 GiB32.00 GiB3.62×16 / 48 / 0

48 of 64 layers cache only a 4,096-token window rather than the full context, on a period of 4. Figures assume the default configuration; --swa-full disables the saving entirely.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 16.77 GiB. The real file is 18.02 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 8.00 GiB at 32K context where the real figure is 2.84 GiB, because most of this model's layers cache a fixed window rather than the whole context.

Architecture

from config.json
Layers
64
Attention heads
40
KV heads
8
Head dim
128
Hidden size
5120
Vocab
102,400
Sliding window
4096
SWA period
4
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does EXAONE-4.0-32B need?
Q4_K_M is exactly 19,343,826,720 bytes (18.02 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is EXAONE-4.0-32B's KV cache?
2.84 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of EXAONE-4.0-32B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.