HuggingFaceTB · text

SmolLM3-3B

HuggingFaceTB/SmolLM3-3B

SmolLM3-3B at Q4_K_M is exactly 1,915,305,312 bytes (1.78 GiB / 1.92 GB) — an effective 4.983 bits per weight, not the nominal 4. Its KV cache at 32K is 2.25 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
3.1B
Architecture
smollm3
36 layers
Context
65,536
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
UD-IQ2_XXS0.89 GiB955,159,0722.485unsloth
IQ2_M1.05 GiB1,125,355,3282.928bartowski
UD-IQ2_M1.07 GiB1,150,607,9042.993unsloth
Q2_K1.17 GiB1,253,302,0803.260bartowski
Q2_K1.17 GiB1,253,302,8163.260unsloth
Q2_K_L1.17 GiB1,253,302,8163.260unsloth
IQ3_XXS1.18 GiB1,267,666,7523.298bartowski
UD-IQ3_XXS1.20 GiB1,285,218,8483.344unsloth
Q2_K_L1.23 GiB1,316,917,0563.426bartowski
IQ3_XS1.28 GiB1,371,414,3363.568bartowski
Q3_K_S1.33 GiB1,432,313,6643.726bartowski
Q3_K_S1.33 GiB1,432,314,4003.726unsloth
IQ3_M1.37 GiB1,469,357,8883.823bartowski
Q3_K_M1.46 GiB1,571,069,7604.087bartowski
Q3_K_M1.46 GiB1,571,070,4964.087unsloth
Q3_K_L1.57 GiB1,690,214,2084.397bartowski
IQ4_XS1.61 GiB1,723,834,1764.485bartowski
IQ4_XS1.61 GiB1,723,834,9124.485unsloth
IQ4_NL1.69 GiB1,810,538,3044.710bartowski
IQ4_NL1.69 GiB1,810,539,0404.710unsloth
Q4_01.69 GiB1,811,455,8084.713bartowski
Q4_01.69 GiB1,811,456,5444.713unsloth
Q4_K_S1.69 GiB1,817,616,1924.729bartowski
Q4_K_S1.69 GiB1,817,616,9284.729unsloth
Q4_K_M1.78 GiB1,915,305,3124.983ggml-org
Q4_K_M1.78 GiB1,915,305,7924.983bartowski
Q4_K_M1.78 GiB1,915,306,5284.983unsloth
Q4_K_L1.84 GiB1,978,920,7685.148bartowski
Q4_11.85 GiB1,981,587,2645.155bartowski
Q4_11.85 GiB1,981,588,0005.155unsloth
Q5_K_S2.01 GiB2,157,354,8165.612bartowski
Q5_K_S2.01 GiB2,157,355,5525.612unsloth
Q5_K_M2.06 GiB2,213,756,7365.759bartowski
Q5_K_M2.06 GiB2,213,757,4725.759unsloth
Q5_K_L2.12 GiB2,277,371,7125.925bartowski
Q6_K2.36 GiB2,530,860,8646.584bartowski
Q6_K2.36 GiB2,530,861,6006.584unsloth
Q6_K_L2.42 GiB2,594,475,8406.750bartowski
Q8_03.05 GiB3,275,574,6248.521ggml-org
Q8_03.05 GiB3,275,575,1048.521bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.28 GiB0.28 GiB36 / 0 / 0
8,1920.56 GiB0.56 GiB36 / 0 / 0
16,3841.13 GiB1.13 GiB36 / 0 / 0
32,7682.25 GiB2.25 GiB36 / 0 / 0
65,5364.50 GiB4.50 GiB36 / 0 / 0
131,0729.00 GiB9.00 GiB36 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 1.61 GiB. The real file is 1.78 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
36
Attention heads
16
KV heads
4
Head dim
128
Hidden size
2048
Vocab
128,256
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window
false

Questions people ask

How much VRAM does SmolLM3-3B need?
Q4_K_M is exactly 1,915,305,312 bytes (1.78 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is SmolLM3-3B's KV cache?
2.25 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of SmolLM3-3B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.