nvidia · text

Llama-3_3-Nemotron-Super-49B-v1_5

nvidia/Llama-3_3-Nemotron-Super-49B-v1_5

Llama-3_3-Nemotron-Super-49B-v1_5 at Q4_K_M is exactly 30,215,578,432 bytes (28.14 GiB / 30.22 GB) — an effective 4.847 bits per weight, not the nominal 4. Its KV cache at 32K is 80.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
49.9B
Architecture
deci
80 layers
Context
131,072
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
UD-IQ1_S10.66 GiB11,446,495,1361.836unsloth
UD-IQ1_M11.52 GiB12,373,305,2481.985unsloth
IQ2_XXS12.72 GiB13,659,170,3042.191bartowski
UD-IQ2_XXS12.99 GiB13,950,298,0162.238unsloth
IQ2_XS14.04 GiB15,076,582,9122.419bartowski
IQ2_S14.76 GiB15,848,482,3042.542bartowski
IQ2_M15.98 GiB17,163,134,4642.753bartowski
UD-IQ2_M16.15 GiB17,341,917,0882.782unsloth
Q2_K17.45 GiB18,739,798,8483.006timteh673
Q2_K17.45 GiB18,739,799,5523.006bartowski
Q2_K17.45 GiB18,739,799,9683.006unsloth
Q2_K_L17.68 GiB18,986,051,4883.046unsloth
IQ3_XXS18.18 GiB19,519,022,5923.131bartowski
UD-IQ3_XXS18.34 GiB19,692,431,2643.159unsloth
Q2_K_L18.41 GiB19,765,847,5523.171bartowski
IQ3_XS19.47 GiB20,908,008,9603.354bartowski
Q3_K_S20.45 GiB21,955,339,7763.522bartowski
Q3_K_S20.45 GiB21,955,340,1923.522unsloth
IQ3_M21.10 GiB22,657,230,3363.635bartowski
Q3_K_M22.64 GiB24,311,227,2003.900timteh673
Q3_K_M22.64 GiB24,311,227,9043.900bartowski
Q3_K_M22.64 GiB24,311,228,3203.900unsloth
Q3_K_L24.47 GiB26,272,065,0244.215bartowski
IQ4_XS25.03 GiB26,871,408,1284.311bartowski
IQ4_XS25.06 GiB26,904,242,0804.316unsloth
IQ4_NL26.43 GiB28,384,044,5444.553bartowski
IQ4_NL26.43 GiB28,384,044,9604.553unsloth
Q4_026.50 GiB28,457,444,8644.565bartowski
Q4_026.50 GiB28,457,445,2804.565unsloth
Q4_K_S26.67 GiB28,633,605,6324.594bartowski
Q4_K_S26.67 GiB28,633,606,0484.594unsloth
Q4_K_M28.14 GiB30,215,578,4324.847timteh673
Q4_K_M28.14 GiB30,215,579,1364.847bartowski
Q4_K_M28.14 GiB30,215,579,5524.847unsloth
Q4_K_L28.87 GiB30,995,375,6164.973bartowski
Q4_129.23 GiB31,383,627,2645.035bartowski
Q4_129.23 GiB31,383,627,6805.035unsloth
Q5_K_S32.07 GiB34,434,590,2085.524bartowski
Q5_K_S32.07 GiB34,434,590,6245.524unsloth
Q5_K_M32.96 GiB35,391,611,7125.678timteh673

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,09610.00 GiB10.00 GiB80 / 0 / 0
8,19220.00 GiB20.00 GiB80 / 0 / 0
16,38440.00 GiB40.00 GiB80 / 0 / 0
32,76880.00 GiB80.00 GiB80 / 0 / 0
65,536160.00 GiB160.00 GiB80 / 0 / 0
131,072320.00 GiB320.00 GiB80 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 26.12 GiB. The real file is 28.14 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
80
Attention heads
64
KV heads
64
Head dim
128
Hidden size
8192
Vocab
128,256
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Llama-3_3-Nemotron-Super-49B-v1_5 need?
Q4_K_M is exactly 30,215,578,432 bytes (28.14 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Llama-3_3-Nemotron-Super-49B-v1_5's KV cache?
80.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Llama-3_3-Nemotron-Super-49B-v1_5 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.