swiss-ai · text

Apertus-8B-Instruct-2509

swiss-ai/Apertus-8B-Instruct-2509

Apertus-8B-Instruct-2509 at Q4_K_M is exactly 5,057,885,440 bytes (4.71 GiB / 5.06 GB) — an effective 5.024 bits per weight, not the nominal 4. Its KV cache at 32K is 4.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
8.1B
Architecture
apertus
32 layers
Context
65,536
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S1.91 GiB2,047,031,0722.034mradermacher
I1-IQ1_M2.04 GiB2,186,622,7522.172mradermacher
UD-IQ1_S2.08 GiB2,230,400,2562.216unsloth
UD-IQ1_M2.19 GiB2,348,758,2722.333unsloth
I1-IQ2_XXS2.25 GiB2,419,275,5522.403mradermacher
UD-IQ2_XXS2.38 GiB2,552,247,5522.535unsloth
I1-IQ2_XS2.44 GiB2,622,175,0082.605mradermacher
I1-IQ2_S2.60 GiB2,787,981,0882.769mradermacher
IQ2_M2.77 GiB2,974,102,9122.954bartowski
I1-IQ2_M2.77 GiB2,974,103,3282.954mradermacher
UD-IQ2_M2.82 GiB3,026,105,6003.006unsloth
I1-Q2_K_S2.82 GiB3,029,677,8563.010mradermacher
Q2_K3.06 GiB3,287,889,1523.266unsloth
Q2_K3.06 GiB3,287,889,2803.266bartowski
IQ3_XXS3.06 GiB3,287,889,2803.266bartowski
I1-Q2_K3.06 GiB3,287,889,6963.266mradermacher
I1-IQ3_XXS3.06 GiB3,287,889,6963.266mradermacher
UD-IQ3_XXS3.13 GiB3,356,538,1123.334unsloth
Q2_K_L3.18 GiB3,413,718,2723.391unsloth
IQ3_XS3.32 GiB3,566,286,2083.543bartowski
I1-IQ3_XS3.32 GiB3,566,286,6243.543mradermacher
Q3_K_S3.43 GiB3,679,008,0003.655unsloth
Q3_K_S3.43 GiB3,679,008,1283.655bartowski
I1-Q3_K_S3.43 GiB3,679,008,5443.655mradermacher
I1-IQ3_S3.44 GiB3,696,834,3363.672mradermacher
Q2_K_L3.55 GiB3,812,177,2803.787bartowski
IQ3_M3.55 GiB3,814,929,7923.790bartowski
I1-IQ3_M3.55 GiB3,814,930,2083.790mradermacher
Q3_K_M3.88 GiB4,165,547,2644.138unsloth
Q3_K_M3.88 GiB4,165,547,3924.138bartowski
I1-Q3_K_M3.88 GiB4,165,547,8084.138mradermacher
IQ4_XS4.16 GiB4,463,342,9764.434bartowski
I1-IQ4_XS4.16 GiB4,463,343,3924.434mradermacher
IQ4_XS4.17 GiB4,480,120,0644.450unsloth
Q3_K_L4.26 GiB4,578,686,3364.548bartowski
I1-Q3_K_L4.26 GiB4,578,686,7524.548mradermacher
IQ4_NL4.37 GiB4,694,029,5684.663unsloth
IQ4_NL4.37 GiB4,694,029,6964.663bartowski
I1-IQ4_NL4.37 GiB4,694,030,1124.663mradermacher
Q4_04.38 GiB4,699,272,4484.668unsloth

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.50 GiB0.50 GiB32 / 0 / 0
8,1921.00 GiB1.00 GiB32 / 0 / 0
16,3842.00 GiB2.00 GiB32 / 0 / 0
32,7684.00 GiB4.00 GiB32 / 0 / 0
65,5368.00 GiB8.00 GiB32 / 0 / 0
131,07216.00 GiB16.00 GiB32 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.22 GiB. The real file is 4.71 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
32
Attention heads
32
KV heads
8
Head dim
128
Hidden size
4096
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Apertus-8B-Instruct-2509 need?
Q4_K_M is exactly 5,057,885,440 bytes (4.71 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Apertus-8B-Instruct-2509's KV cache?
4.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Apertus-8B-Instruct-2509 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.