swiss-ai · text

Apertus-70B-Instruct-2509

swiss-ai/Apertus-70B-Instruct-2509

Apertus-70B-Instruct-2509 at Q4_K_M is exactly 43,721,584,608 bytes (40.72 GiB / 43.72 GB) — an effective 4.954 bits per weight, not the nominal 4. Its KV cache at 32K is 10.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
70.6B
Architecture
apertus
80 layers
Context
65,536
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
UD-IQ1_S15.46 GiB16,599,131,8081.881unsloth
IQ1_M15.74 GiB16,903,218,9761.915bartowski
UD-IQ1_M16.58 GiB17,799,227,0402.017unsloth
IQ2_XXS17.88 GiB19,203,532,5762.176bartowski
UD-IQ2_XXS18.54 GiB19,908,437,6642.256unsloth
IQ2_XS19.75 GiB21,211,555,6162.404bartowski
IQ2_S20.89 GiB22,433,408,8002.542bartowski
IQ2_M22.61 GiB24,273,659,6802.751bartowski
UD-IQ2_M23.12 GiB24,825,341,6002.813unsloth
Q2_K25.40 GiB27,272,048,6083.090giladgd
Q2_K25.40 GiB27,272,062,6243.090unsloth
Q2_K25.40 GiB27,272,062,7523.090bartowski
IQ3_XXS25.53 GiB27,411,523,3603.106bartowski
Q2_K_L25.63 GiB27,523,720,8643.119unsloth
UD-IQ3_XXS26.01 GiB27,932,665,5043.165unsloth
Q2_K_L26.38 GiB28,320,638,7523.209bartowski
IQ3_XS27.55 GiB29,583,124,2563.352bartowski
Q3_K_S28.65 GiB30,768,000,9923.486giladgd
Q3_K_S28.65 GiB30,768,015,0083.486unsloth
Q3_K_S28.65 GiB30,768,015,1363.486bartowski
IQ3_M29.84 GiB32,038,102,8163.630bartowski
Q3_K_M33.10 GiB35,535,876,0644.027giladgd
Q3_K_M33.10 GiB35,535,890,0804.027unsloth
Q3_K_M33.10 GiB35,535,890,2084.027bartowski
IQ4_XS35.33 GiB37,933,983,5204.298bartowski
IQ4_XS35.36 GiB37,967,537,8244.302unsloth
Q3_K_L36.87 GiB39,591,768,0324.486giladgd
Q3_K_L36.87 GiB39,591,782,1764.486bartowski
Q4_037.25 GiB40,001,761,2484.533giladgd
IQ4_NL37.33 GiB40,085,661,3444.542unsloth
IQ4_NL37.33 GiB40,085,661,4724.542bartowski
Q4_037.46 GiB40,221,976,2244.558unsloth
Q4_037.46 GiB40,221,976,3524.558bartowski
Q4_K_S37.67 GiB40,446,357,4724.583giladgd
Q4_K_S37.67 GiB40,446,371,4884.583unsloth
Q4_K_S37.67 GiB40,446,371,6164.583bartowski
Q4_K_M40.72 GiB43,721,584,6084.954giladgd
Q4_K_M40.72 GiB43,721,598,6244.954unsloth
Q4_K_M40.72 GiB43,721,598,7524.954bartowski
Q4_141.30 GiB44,347,074,2085.025unsloth

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.25 GiB1.25 GiB80 / 0 / 0
8,1922.50 GiB2.50 GiB80 / 0 / 0
16,3845.00 GiB5.00 GiB80 / 0 / 0
32,76810.00 GiB10.00 GiB80 / 0 / 0
65,53620.00 GiB20.00 GiB80 / 0 / 0
131,07240.00 GiB40.00 GiB80 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 36.99 GiB. The real file is 40.72 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
80
Attention heads
64
KV heads
8
Head dim
128
Hidden size
8192
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Apertus-70B-Instruct-2509 need?
Q4_K_M is exactly 43,721,584,608 bytes (40.72 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Apertus-70B-Instruct-2509's KV cache?
10.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Apertus-70B-Instruct-2509 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.