nvidia · text

Llama-3_3-Nemotron-Super-49B-v1

nvidia/Llama-3_3-Nemotron-Super-49B-v1

Llama-3_3-Nemotron-Super-49B-v1 at Q4_K_M is exactly 30,215,576,224 bytes (28.14 GiB / 30.22 GB) — an effective 4.847 bits per weight, not the nominal 4. Its KV cache at 32K is 80.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
49.9B
Architecture
deci
80 layers
Context
131,072
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S10.27 GiB11,029,863,3281.770mmnga
UD-IQ1_S10.66 GiB11,446,492,6721.836unsloth
IQ1_M11.19 GiB12,015,852,4481.928mmnga
UD-IQ1_M11.52 GiB12,373,302,7841.985unsloth
IQ2_XXS12.72 GiB13,659,167,6482.191mmnga
IQ2_XXS12.72 GiB13,659,167,6802.191bartowski
UD-IQ2_XXS12.99 GiB13,950,295,5522.238unsloth
IQ2_XS14.04 GiB15,076,580,2562.419mmnga
IQ2_XS14.04 GiB15,076,580,2882.419bartowski
IQ2_S14.76 GiB15,848,479,6482.542mmnga
IQ2_S14.76 GiB15,848,479,6802.542bartowski
IQ2_M15.98 GiB17,163,131,8082.753mmnga
IQ2_M15.98 GiB17,163,131,8402.753bartowski
UD-IQ2_M16.15 GiB17,341,914,6242.782unsloth
Q2_K17.45 GiB18,739,796,6403.006mmnga
Q2_K17.45 GiB18,739,796,9283.006bartowski
Q2_K17.45 GiB18,739,797,5043.006unsloth
Q2_K_L17.68 GiB18,986,049,0243.046unsloth
IQ3_XXS18.18 GiB19,519,019,9363.131mmnga
IQ3_XXS18.18 GiB19,519,019,9683.131bartowski
UD-IQ3_XXS18.34 GiB19,692,428,8003.159unsloth
Q2_K_L18.41 GiB19,765,844,9283.171bartowski
IQ3_XS19.47 GiB20,908,006,3043.354mmnga
IQ3_XS19.47 GiB20,908,006,3363.354bartowski
Q3_K_S20.45 GiB21,955,336,8643.522mmnga
IQ3_S20.45 GiB21,955,337,1203.522mmnga
Q3_K_S20.45 GiB21,955,337,1523.522bartowski
Q3_K_S20.45 GiB21,955,337,7283.522unsloth
IQ3_M21.10 GiB22,657,227,6803.635mmnga
IQ3_M21.10 GiB22,657,227,7123.635bartowski
Q3_K_M22.64 GiB24,311,224,9923.900mmnga
Q3_K_M22.64 GiB24,311,225,2803.900bartowski
Q3_K_M22.64 GiB24,311,225,8563.900unsloth
Q3_K_L24.47 GiB26,272,062,1124.215mmnga
Q3_K_L24.47 GiB26,272,062,4004.215bartowski
IQ4_XS25.03 GiB26,871,405,4724.311mmnga
IQ4_XS25.03 GiB26,871,405,5044.311bartowski
IQ4_XS25.06 GiB26,904,239,6164.316unsloth
Q4_026.39 GiB28,332,661,4084.545mmnga
IQ4_NL26.43 GiB28,384,041,8884.553mmnga

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,09610.00 GiB10.00 GiB80 / 0 / 0
8,19220.00 GiB20.00 GiB80 / 0 / 0
16,38440.00 GiB40.00 GiB80 / 0 / 0
32,76880.00 GiB80.00 GiB80 / 0 / 0
65,536160.00 GiB160.00 GiB80 / 0 / 0
131,072320.00 GiB320.00 GiB80 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 26.12 GiB. The real file is 28.14 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
80
Attention heads
64
KV heads
64
Head dim
128
Hidden size
8192
Vocab
128,256
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Llama-3_3-Nemotron-Super-49B-v1 need?
Q4_K_M is exactly 30,215,576,224 bytes (28.14 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Llama-3_3-Nemotron-Super-49B-v1's KV cache?
80.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Llama-3_3-Nemotron-Super-49B-v1 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.