nvidia · text

Llama-3_1-Nemotron-51B-Instruct

nvidia/Llama-3_1-Nemotron-51B-Instruct

Llama-3_1-Nemotron-51B-Instruct at Q4_K_M is exactly 31,037,307,136 bytes (28.91 GiB / 31.04 GB) — an effective 4.821 bits per weight, not the nominal 4. Its KV cache at 32K is 80.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
51.5B
Architecture
deci
80 layers
Context
131,072
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XXS13.07 GiB14,037,244,1602.180bartowski
IQ2_XS14.46 GiB15,522,912,5122.411bartowski
IQ2_S15.33 GiB16,458,193,1522.557bartowski
IQ2_M16.57 GiB17,792,866,5602.764bartowski
Q2_K18.09 GiB19,418,642,6883.016bartowski
Q2_K_L19.04 GiB20,444,690,6883.176bartowski
IQ3_XS20.09 GiB21,566,970,1123.350bartowski
Q3_K_S21.10 GiB22,652,393,7283.519bartowski
IQ3_M21.88 GiB23,489,091,8403.649bartowski
Q3_K_M23.45 GiB25,182,345,4723.912bartowski
Q3_K_L25.47 GiB27,349,752,0644.248bartowski
IQ4_XS25.83 GiB27,736,619,2644.309bartowski
IQ4_NL27.29 GiB29,300,996,3524.551bartowski
Q4_027.33 GiB29,344,119,0404.558bartowski
Q4_K_S27.46 GiB29,484,497,1524.580bartowski
Q4_K_M28.91 GiB31,037,307,1364.821bartowski
Q4_K_L29.63 GiB31,817,103,6164.942bartowski
Q4_130.18 GiB32,405,436,6725.034bartowski
Q5_K_S33.12 GiB35,558,504,7045.524bartowski
Q5_K_M33.96 GiB36,465,391,8725.664bartowski
Q5_K_L34.56 GiB37,113,854,2085.765bartowski
Q6_K39.36 GiB42,258,774,2726.564bartowski
Q6_K_L39.83 GiB42,767,694,0806.643bartowski
Q8_02 shards50.97 GiB54,731,373,0248.502bartowski
F163 shards95.94 GiB103,012,399,39216.002bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,09610.00 GiB10.00 GiB80 / 0 / 0
8,19220.00 GiB20.00 GiB80 / 0 / 0
16,38440.00 GiB40.00 GiB80 / 0 / 0
32,76880.00 GiB80.00 GiB80 / 0 / 0
65,536160.00 GiB160.00 GiB80 / 0 / 0
131,072320.00 GiB320.00 GiB80 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 26.98 GiB. The real file is 28.91 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
80
Attention heads
64
KV heads
64
Head dim
128
Hidden size
8192
Vocab
128,256
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Llama-3_1-Nemotron-51B-Instruct need?
Q4_K_M is exactly 31,037,307,136 bytes (28.91 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Llama-3_1-Nemotron-51B-Instruct's KV cache?
80.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Llama-3_1-Nemotron-51B-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.