LGAI-EXAONE · text

EXAONE-Deep-7.8B

LGAI-EXAONE/EXAONE-Deep-7.8B

EXAONE-Deep-7.8B at Q4_K_M is exactly 4,770,649,856 bytes (4.44 GiB / 4.77 GB) — an effective 4.881 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
7.8B
Architecture
exaone
null layers
Context
32,768
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M2.63 GiB2,826,328,6722.892bartowski
Q2_K2.84 GiB3,053,869,6643.125bartowski
IQ3_XXS2.94 GiB3,152,960,0963.226bartowski
IQ3_XS3.15 GiB3,382,729,3123.461bartowski
Q2_K_L3.23 GiB3,463,469,6643.544bartowski
Q3_K_S3.29 GiB3,528,481,3763.610bartowski
IQ3_M3.40 GiB3,648,805,4723.733bartowski
Q3_K_M3.62 GiB3,882,900,0643.973bartowski
Q3_K_L3.90 GiB4,185,938,5284.283bartowski
IQ4_XS4.01 GiB4,300,888,6724.401bartowski
IQ4_XS4.04 GiB4,337,587,9684.438LGAI-EXAONE
Q4_04.21 GiB4,525,808,2244.631bartowski
IQ4_NL4.22 GiB4,527,905,3764.633bartowski
Q4_K_S4.23 GiB4,542,585,4404.648bartowski
Q4_K_M4.44 GiB4,770,649,8564.881LGAI-EXAONE
Q4_K_M4.44 GiB4,770,650,7204.881bartowski
Q4_14.63 GiB4,973,550,1765.089bartowski
Q4_K_L4.73 GiB5,081,946,7205.200bartowski
Q5_K_S5.06 GiB5,435,972,1925.562bartowski
Q5_K_M5.19 GiB5,569,664,7685.699LGAI-EXAONE
Q5_K_M5.19 GiB5,569,665,6325.699bartowski
Q5_K_L5.43 GiB5,828,532,8325.964bartowski
Q6_K5.98 GiB6,418,618,1126.568LGAI-EXAONE
Q6_K5.98 GiB6,418,618,9766.568bartowski
Q6_K_L6.17 GiB6,621,780,5766.776bartowski
Q8_07.74 GiB8,312,084,2248.505LGAI-EXAONE
Q8_07.74 GiB8,312,085,0888.505bartowski
BF1614.57 GiB15,641,630,46416.005LGAI-EXAONE
BF1614.57 GiB15,641,631,04016.005bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.10 GiB. The real file is 4.44 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
Attention heads
32
KV heads
8
Head dim
128
Hidden size
4096
Vocab
102,400
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does EXAONE-Deep-7.8B need?
Q4_K_M is exactly 4,770,649,856 bytes (4.44 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of EXAONE-Deep-7.8B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.