mistralai · text

Devstral-2-123B-Instruct-2512

mistralai/Devstral-2-123B-Instruct-2512

Devstral-2-123B-Instruct-2512 at Q4_K_M is exactly 74,897,652,896 bytes (69.75 GiB / 74.90 GB) — an effective 4.793 bits per weight, not the nominal 4. Its KV cache at 32K is 11.00 GiB.

From the file· summed from 2 file(s)From the file· KV per layer
Parameters
125B
Architecture
llama
88 layers
Context
262,144
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S25.33 GiB27,194,260,8321.740bartowski
UD-IQ1_S26.49 GiB28,445,876,4481.820unsloth
IQ1_M27.59 GiB29,620,796,7681.895bartowski
UD-IQ1_M28.47 GiB30,565,310,6881.956unsloth
IQ2_XXS31.35 GiB33,665,023,3282.154bartowski
UD-IQ2_XXS32.04 GiB34,404,671,7122.201unsloth
IQ2_XS34.75 GiB37,315,640,6722.388bartowski
IQ2_S37.01 GiB39,741,390,1762.543bartowski
IQ2_M40.03 GiB42,976,771,4242.750bartowski
UD-IQ2_M40.55 GiB43,543,404,7682.786unsloth
Q2_K43.39 GiB46,591,212,8962.981bartowski
Q2_K43.39 GiB46,591,221,9842.981unsloth
Q2_K_L43.74 GiB46,968,709,3443.005unsloth
Q2_K_L44.86 GiB48,164,076,8963.082bartowski
IQ3_XXS45.04 GiB48,366,189,9203.095bartowski
UD-IQ3_XXS45.60 GiB48,963,887,3283.133unsloth
IQ3_XS2 shards48.11 GiB51,659,767,3603.305bartowski
Q3_K_S2 shards50.63 GiB54,367,452,7043.479bartowski
Q3_K_S2 shards50.63 GiB54,367,461,7923.479unsloth
IQ3_M2 shards52.89 GiB56,793,988,6723.634bartowski
Q3_K_M2 shards56.46 GiB60,620,373,5683.879bartowski
Q3_K_M2 shards56.46 GiB60,620,382,6563.879unsloth
Q3_K_L2 shards61.53 GiB66,071,919,7764.228lmstudio-community
Q3_K_L2 shards61.53 GiB66,071,920,1604.228bartowski
IQ4_XS2 shards62.47 GiB67,074,620,9604.292bartowski
IQ4_XS2 shards62.51 GiB67,124,961,6964.295unsloth
IQ4_NL2 shards66.03 GiB70,896,680,4804.536bartowski
IQ4_NL2 shards66.03 GiB70,896,689,5684.536unsloth
Q4_02 shards66.12 GiB71,000,489,5364.543bartowski
Q4_02 shards66.12 GiB71,000,498,5924.543unsloth
Q4_K_S2 shards66.36 GiB71,249,002,0164.559bartowski
Q4_K_S2 shards66.36 GiB71,249,011,1044.559unsloth
Q4_K_M2 shards69.75 GiB74,897,652,8964.793lmstudio-community
Q4_K_M2 shards69.75 GiB74,897,653,3124.793bartowski
Q4_K_M2 shards69.75 GiB74,897,662,4004.793unsloth
Q4_K_L2 shards70.87 GiB76,093,029,9204.869bartowski
Q4_12 shards73.08 GiB78,471,593,5045.021bartowski
Q4_12 shards73.08 GiB78,471,602,5925.021unsloth
Q5_K_S3 shards80.27 GiB86,184,918,6885.515bartowski
Q5_K_S2 shards80.27 GiB86,184,927,6805.515unsloth

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.38 GiB1.38 GiB88 / 0 / 0
8,1922.75 GiB2.75 GiB88 / 0 / 0
16,3845.50 GiB5.50 GiB88 / 0 / 0
32,76811.00 GiB11.00 GiB88 / 0 / 0
65,53622.00 GiB22.00 GiB88 / 0 / 0
131,07244.00 GiB44.00 GiB88 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 65.50 GiB. The real file is 69.75 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
88
Attention heads
96
KV heads
8
Head dim
128
Hidden size
12288
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Devstral-2-123B-Instruct-2512 need?
Q4_K_M is exactly 74,897,652,896 bytes (69.75 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Devstral-2-123B-Instruct-2512's KV cache?
11.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Devstral-2-123B-Instruct-2512 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.