mistralai · text

Mistral-Large-Instruct-2407

mistralai/Mistral-Large-Instruct-2407

Mistral-Large-Instruct-2407 at Q4_K_M is exactly 73,219,622,816 bytes (68.19 GiB / 73.22 GB) — an effective 4.777 bits per weight, not the nominal 4.

From the file· summed from 2 file(s)
Parameters
123B
Architecture
llama
Context
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S24.18 GiB25,959,777,9841.694InferenceIllusionist
IQ1_S24.18 GiB25,959,778,0481.694MaziyarPanahi
IQ1_M26.44 GiB28,386,313,9201.852InferenceIllusionist
IQ1_M26.44 GiB28,386,313,9841.852MaziyarPanahi
IQ1_M26.44 GiB28,386,317,9841.852bartowski
IQ2_XXS30.20 GiB32,430,540,4802.116InferenceIllusionist
IQ2_XXS30.20 GiB32,430,544,5442.116bartowski
IQ2_XS33.60 GiB36,081,157,8242.354InferenceIllusionist
IQ2_XS33.60 GiB36,081,157,8882.354MaziyarPanahi
IQ2_XS33.60 GiB36,081,161,8882.354bartowski
IQ2_S35.75 GiB38,384,223,9362.505InferenceIllusionist
IQ2_M38.76 GiB41,619,605,1842.716InferenceIllusionist
IQ2_M38.76 GiB41,619,609,2482.716bartowski
Q2_K42.09 GiB45,196,297,9202.949InferenceIllusionist
Q2_K42.09 GiB45,196,297,9842.949MaziyarPanahi
Q2_K42.09 GiB45,196,301,9842.949bartowski
Q2_K_L42.46 GiB45,589,517,9842.975bartowski
IQ3_XXS43.78 GiB47,009,023,6803.067InferenceIllusionist
IQ3_XXS43.78 GiB47,009,027,7443.067bartowski
IQ3_XS2 shards46.70 GiB50,142,168,9603.272InferenceIllusionist
Q3_K_S2 shards49.22 GiB52,849,854,3683.448InferenceIllusionist
Q3_K_S2 shards49.22 GiB52,849,858,4323.448bartowski
IQ3_S2 shards49.36 GiB52,996,917,1523.458InferenceIllusionist
IQ3_M2 shards51.48 GiB55,276,390,3043.607InferenceIllusionist
IQ3_M2 shards51.48 GiB55,276,394,3683.607bartowski
Q3_K_M2 shards55.04 GiB59,102,775,1683.856InferenceIllusionist
Q3_K_M2 shards55.04 GiB59,102,779,2643.856bartowski
Q3_K_L2 shards60.12 GiB64,554,325,8884.212bartowski
IQ4_XS2 shards60.94 GiB65,434,339,2004.269InferenceIllusionist
IQ4_XS2 shards60.94 GiB65,434,343,2644.269bartowski
Q4_02 shards64.56 GiB69,322,463,1044.523bartowski
Q4_K_S2 shards64.79 GiB69,570,971,5524.539InferenceIllusionist
Q4_K_S2 shards64.79 GiB69,570,975,6164.539bartowski
Q4_K_M2 shards68.19 GiB73,219,622,8164.777InferenceIllusionist
Q4_K_M2 shards68.19 GiB73,219,626,8804.777bartowski
Q5_K_S2 shards78.56 GiB84,355,893,1205.504InferenceIllusionist
Q5_K_S3 shards78.56 GiB84,355,894,4325.504bartowski
Q5_K_M2 shards80.55 GiB86,488,303,5205.643InferenceIllusionist
Q5_K_M3 shards80.55 GiB86,488,307,6805.643bartowski
Q6_K3 shards93.68 GiB100,586,276,8646.563InferenceIllusionist

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 64.23 GiB. The real file is 68.19 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Mistral-Large-Instruct-2407 need?
Q4_K_M is exactly 73,219,622,816 bytes (68.19 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Mistral-Large-Instruct-2407 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.