LGAI-EXAONE · text

EXAONE-3.5-7.8B-Instruct

LGAI-EXAONE/EXAONE-3.5-7.8B-Instruct

EXAONE-3.5-7.8B-Instruct at Q4_K_M is exactly 4,770,650,016 bytes (4.44 GiB / 4.77 GB) — an effective 4.881 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
7.8B
Architecture
exaone
null layers
Context
32,768
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_S2.46 GiB2,636,536,0002.698bartowski
IQ2_M2.63 GiB2,826,328,2562.892bartowski
Q2_K2.84 GiB3,053,869,2483.125bartowski
IQ3_XS3.15 GiB3,382,728,8963.461bartowski
Q2_K_L3.23 GiB3,463,469,2483.544bartowski
Q3_K_S3.29 GiB3,528,480,9603.610bartowski
IQ3_M3.40 GiB3,648,805,0563.733bartowski
Q3_K_M3.62 GiB3,882,899,6483.973bartowski
Q3_K_L3.90 GiB4,185,937,8244.283lmstudio-community
Q3_K_L3.90 GiB4,185,938,1124.283bartowski
IQ4_XS4.01 GiB4,300,888,2564.401bartowski
Q4_04.21 GiB4,525,807,8084.631bartowski
IQ4_NL4.22 GiB4,527,904,9604.633bartowski
Q4_K_S4.23 GiB4,542,585,0244.648bartowski
Q4_K_M4.44 GiB4,770,650,0164.881lmstudio-community
Q4_K_M4.44 GiB4,770,650,3044.881bartowski
Q4_K_L4.73 GiB5,081,946,3045.200bartowski
Q5_K_S5.06 GiB5,435,971,7765.562bartowski
Q5_K_M5.19 GiB5,569,665,2165.699bartowski
Q5_K_L5.43 GiB5,828,532,4165.964bartowski
Q6_K5.98 GiB6,418,618,2726.568lmstudio-community
Q6_K5.98 GiB6,418,618,5606.568bartowski
Q6_K_L6.17 GiB6,621,780,1606.776bartowski
Q8_07.74 GiB8,312,084,3848.505lmstudio-community
Q8_07.74 GiB8,312,084,6728.505bartowski
F1614.57 GiB15,641,630,91216.005bartowski
F3229.13 GiB31,277,995,93632.004bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.10 GiB. The real file is 4.44 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
Attention heads
32
KV heads
8
Head dim
128
Hidden size
4096
Vocab
102,400
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does EXAONE-3.5-7.8B-Instruct need?
Q4_K_M is exactly 4,770,650,016 bytes (4.44 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of EXAONE-3.5-7.8B-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.