HuggingFaceTB · text

SmolLM-1.7B-Instruct-v0.2

HuggingFaceTB/SmolLM-1.7B-Instruct-v0.2

SmolLM-1.7B-Instruct-v0.2 at Q4_K_M is exactly 1,055,610,144 bytes (0.98 GiB / 1.06 GB) — an effective 4.935 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
1.7B
Architecture
llama
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S0.38 GiB412,308,7681.927MaziyarPanahi
IQ1_M0.41 GiB444,159,2642.076MaziyarPanahi
IQ2_XS0.51 GiB542,856,4802.538MaziyarPanahi
Q2_K0.63 GiB674,583,8403.153MaziyarPanahi
Q2_K0.63 GiB674,583,9043.153bartowski
IQ3_XXS0.63 GiB680,088,9283.179bartowski
Q2_K_L0.65 GiB698,963,2963.267bartowski
IQ3_XS0.69 GiB739,071,2643.455MaziyarPanahi
IQ3_XS0.69 GiB739,071,3283.455bartowski
Q3_K_S0.72 GiB776,820,0003.631MaziyarPanahi
IQ3_M0.75 GiB810,243,4243.788bartowski
Q3_K_M0.80 GiB860,181,7924.021MaziyarPanahi
Q3_K_L0.87 GiB932,533,5364.359MaziyarPanahi
Q3_K_L0.87 GiB932,533,6004.359bartowski
IQ4_XS0.88 GiB940,397,8564.396MaziyarPanahi
IQ4_XS0.88 GiB940,397,9204.396bartowski
Q4_K_S0.93 GiB999,118,1124.670MaziyarPanahi
Q4_K_S0.93 GiB999,118,1764.670bartowski
Q4_K_M0.98 GiB1,055,610,1444.935MaziyarPanahi
Q4_K_M0.98 GiB1,055,610,2084.935bartowski
Q4_K_L1.01 GiB1,079,989,6005.048bartowski
Q5_K_S1.11 GiB1,192,056,0965.572MaziyarPanahi
Q5_K_S1.11 GiB1,192,056,1605.572bartowski
Q5_K_M1.14 GiB1,225,479,4565.729MaziyarPanahi
Q5_K_M1.14 GiB1,225,479,5205.729bartowski
Q5_K_L1.16 GiB1,249,858,9125.843bartowski
Q6_K1.31 GiB1,405,965,6006.572MaziyarPanahi
Q6_K1.31 GiB1,405,965,6646.572bartowski
Q6_K_L1.33 GiB1,430,345,0566.686bartowski
Q8_01.70 GiB1,820,415,2648.510MaziyarPanahi
Q8_01.70 GiB1,820,415,3288.510bartowski
F326.38 GiB6,847,288,38432.008bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 0.90 GiB. The real file is 0.98 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does SmolLM-1.7B-Instruct-v0.2 need?
Q4_K_M is exactly 1,055,610,144 bytes (0.98 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of SmolLM-1.7B-Instruct-v0.2 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.