nvidia · text

NVIDIA-Nemotron-Nano-9B-v2

nvidia/NVIDIA-Nemotron-Nano-9B-v2

NVIDIA-Nemotron-Nano-9B-v2 at Q4_K_M is exactly 6,525,628,992 bytes (6.08 GiB / 6.53 GB) — an effective 5.873 bits per weight, not the nominal 4. Its KV cache at 32K is 7.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
8.9B
Architecture
nemotron_h
56 layers
Context
131,072
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_S4.62 GiB4,957,808,4804.462bartowski
IQ2_M4.65 GiB4,996,515,6804.497bartowski
Q2_K4.66 GiB5,006,192,1924.506dominguesm
Q2_K4.66 GiB5,006,192,4804.506bartowski
IQ3_XXS4.73 GiB5,073,930,0804.567bartowski
IQ3_XS4.78 GiB5,131,990,8804.619bartowski
Q3_K_S4.78 GiB5,131,990,8804.619bartowski
IQ3_M4.85 GiB5,207,935,8404.688bartowski
IQ4_XS4.91 GiB5,267,107,6804.741bartowski
Q2_K_L4.94 GiB5,299,793,7604.770bartowski
Q4_04.94 GiB5,308,681,7924.778dominguesm
IQ4_NL4.94 GiB5,308,682,0804.778bartowski
Q4_04.97 GiB5,339,414,8804.806bartowski
Q3_K_M5.01 GiB5,379,734,8804.842bartowski
Q3_K_L5.11 GiB5,488,365,9204.940bartowski
Q4_15.43 GiB5,827,358,2725.245dominguesm
Q4_15.43 GiB5,827,358,5605.245bartowski
Q4_K_S5.79 GiB6,211,616,8325.591dominguesm
Q4_K_S5.79 GiB6,211,617,1205.591bartowski
Q5_05.91 GiB6,346,034,7525.712dominguesm
Q4_K_M6.08 GiB6,525,628,9925.873dominguesm
Q4_K_M6.08 GiB6,525,629,2805.873bartowski
Q4_K_L6.28 GiB6,745,830,2406.072bartowski
Q5_K_S6.32 GiB6,781,562,7206.104bartowski
Q5_K_M6.58 GiB7,069,805,6326.363dominguesm
Q5_K_M6.58 GiB7,069,805,9206.363bartowski
Q5_K_L6.76 GiB7,253,306,7206.529bartowski
Q6_K8.51 GiB9,135,892,0328.223dominguesm
Q6_K_L8.51 GiB9,135,892,3208.223bartowski
Q6_K8.51 GiB9,135,892,3208.223bartowski
Q8_08.81 GiB9,458,093,6328.513dominguesm
Q8_08.81 GiB9,458,093,9208.513bartowski
BF1616.57 GiB17,788,743,23216.011bartowski
F1616.57 GiB17,788,743,23216.011dominguesm

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.88 GiB0.88 GiB56 / 0 / 0
8,1921.75 GiB1.75 GiB56 / 0 / 0
16,3843.50 GiB3.50 GiB56 / 0 / 0
32,7687.00 GiB7.00 GiB56 / 0 / 0
65,53614.00 GiB14.00 GiB56 / 0 / 0
131,07228.00 GiB28.00 GiB56 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.66 GiB. The real file is 6.08 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
56
Attention heads
40
KV heads
8
Head dim
128
Hidden size
4480
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does NVIDIA-Nemotron-Nano-9B-v2 need?
Q4_K_M is exactly 6,525,628,992 bytes (6.08 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is NVIDIA-Nemotron-Nano-9B-v2's KV cache?
7.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of NVIDIA-Nemotron-Nano-9B-v2 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.