tiiuae · text

Falcon3-1B-Instruct

tiiuae/Falcon3-1B-Instruct

Falcon3-1B-Instruct at Q4_K_M is exactly 1,057,044,608 bytes (0.98 GiB / 1.06 GB) — an effective 5.066 bits per weight, not the nominal 4. Its KV cache at 32K is 2.25 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
1.7B
Architecture
llama
18 layers
Context
8,192
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K0.21 GiB224,950,4001.078tiiuae
IQ2_M0.64 GiB683,735,2003.276bartowski
Q2_K0.68 GiB727,087,2643.484bartowski
IQ3_XS0.75 GiB801,437,8563.841bartowski
Q3_K_S0.77 GiB827,193,5043.964bartowski
IQ3_M0.79 GiB846,690,4644.057bartowski
Q3_K_M0.82 GiB884,963,4564.241tiiuae
Q3_K_M0.82 GiB884,963,4884.241bartowski
Q3_K_L0.87 GiB934,246,5604.477bartowski
IQ4_XS0.90 GiB969,472,1604.646bartowski
Q2_K_L0.92 GiB989,231,2644.740bartowski
Q4_00.94 GiB1,013,250,1764.856tiiuae
IQ4_NL0.94 GiB1,013,250,2084.856bartowski
Q4_00.95 GiB1,015,347,3604.866bartowski
Q4_K_S0.95 GiB1,018,493,0884.881bartowski
Q4_K_M0.98 GiB1,057,044,6085.066tiiuae
Q4_K_M0.98 GiB1,057,044,6405.066bartowski
Q5_01.11 GiB1,188,362,3685.695tiiuae
Q5_K_S1.11 GiB1,188,362,4005.695bartowski
Q5_K_M1.13 GiB1,210,923,1365.803tiiuae
Q5_K_M1.13 GiB1,210,923,1685.803bartowski
Q4_K_L1.17 GiB1,256,274,0806.020bartowski
Q6_K1.28 GiB1,374,419,0726.586tiiuae
Q6_K1.28 GiB1,374,419,1046.586bartowski
Q5_K_L1.28 GiB1,376,598,1766.597bartowski
Q6_K_L1.40 GiB1,504,442,5287.210bartowski
Q8_01.66 GiB1,778,710,6568.524tiiuae
Q8_01.66 GiB1,778,710,6888.524bartowski
F163.11 GiB3,343,710,08016.023bartowski
F163.11 GiB3,343,710,33616.023tiiuae

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.28 GiB0.28 GiB18 / 0 / 0
8,1920.56 GiB0.56 GiB18 / 0 / 0
16,3841.13 GiB1.13 GiB18 / 0 / 0
32,7682.25 GiB2.25 GiB18 / 0 / 0
65,5364.50 GiB4.50 GiB18 / 0 / 0
131,0729.00 GiB9.00 GiB18 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 0.87 GiB. The real file is 0.98 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
18
Attention heads
8
KV heads
4
Head dim
256
Hidden size
2048
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Falcon3-1B-Instruct need?
Q4_K_M is exactly 1,057,044,608 bytes (0.98 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Falcon3-1B-Instruct's KV cache?
2.25 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Falcon3-1B-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.