nvidia · text

Nemotron-Mini-4B-Instruct

nvidia/Nemotron-Mini-4B-Instruct

Nemotron-Mini-4B-Instruct at Q4_K_M is exactly 2,697,387,072 bytes (2.51 GiB / 2.70 GB) — an effective 5.149 bits per weight, not the nominal 4. Its KV cache at 32K is 4.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
4.2B
Architecture
nemotron
32 layers
Context
4,096
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ3_M2.03 GiB2,184,092,7364.170bartowski
Q3_K_L2.28 GiB2,452,954,1764.683bartowski
IQ4_XS2.29 GiB2,461,260,8644.699bartowski
Q4_02.40 GiB2,574,703,6804.915bartowski
Q4_K_S2.41 GiB2,583,354,4324.932bartowski
Q4_K_M2.51 GiB2,697,387,0725.149bartowski
Q5_K_S2.79 GiB2,993,085,5045.714bartowski
Q5_K_M2.85 GiB3,059,932,2245.842bartowski
Q4_K_L3.06 GiB3,281,067,0726.264bartowski
Q6_K3.21 GiB3,445,136,4486.577bartowski
Q5_K_L3.30 GiB3,545,308,2246.768bartowski
Q6_K_L3.56 GiB3,826,064,4487.304bartowski
Q8_04.15 GiB4,459,928,6408.514bartowski
Q2_K2 shards6.43 GiB6,908,987,74413.190second-state
Q3_K_S2 shards6.75 GiB7,247,565,15213.836second-state
Q3_K_M2 shards7.15 GiB7,676,974,94414.656second-state
Q4_02 shards7.34 GiB7,876,307,29615.037second-state
Q3_K_L2 shards7.40 GiB7,941,319,52015.161second-state
F167.81 GiB8,388,156,19216.014bartowski
Q4_K_S2 shards8.19 GiB8,794,970,97616.790second-state
Q4_K_M2 shards8.59 GiB9,223,015,77617.607second-state
Q5_02 shards8.70 GiB9,339,119,96817.829second-state
Q5_K_S2 shards9.10 GiB9,774,647,64818.660second-state
Q5_K_M2 shards9.43 GiB10,129,737,56819.338second-state
Q6_K2 shards11.72 GiB12,581,028,19224.018second-state
Q8_02 shards12.96 GiB13,918,021,98426.571second-state
F162 shards24.38 GiB26,176,899,424second-state

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.50 GiB0.50 GiB32 / 0 / 0
8,1921.00 GiB1.00 GiB32 / 0 / 0
16,3842.00 GiB2.00 GiB32 / 0 / 0
32,7684.00 GiB4.00 GiB32 / 0 / 0
65,5368.00 GiB8.00 GiB32 / 0 / 0
131,07216.00 GiB16.00 GiB32 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 2.20 GiB. The real file is 2.51 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
32
Attention heads
24
KV heads
8
Head dim
128
Hidden size
3072
Vocab
256,000
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Nemotron-Mini-4B-Instruct need?
Q4_K_M is exactly 2,697,387,072 bytes (2.51 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Nemotron-Mini-4B-Instruct's KV cache?
4.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Nemotron-Mini-4B-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.