nvidia · text · mixture of experts

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16

nvidia/NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16

NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 at Q4_K_M is exactly 51,601,807,488 bytes (48.06 GiB / 51.60 GB) — an effective 5.478 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
75.4B
total, not active
Architecture
nemotron_h_moe
null layers
Context
262,144
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K29.30 GiB31,461,152,8963.340RemySkye
Q3_K_S33.22 GiB35,670,087,8083.787RemySkye
Q3_K_M37.78 GiB40,567,117,9524.307RemySkye
IQ4_XS40.00 GiB42,954,905,7284.560RemySkye
Q3_K_L41.18 GiB44,217,473,1524.694RemySkye
Q4_041.43 GiB44,488,284,2884.723RemySkye
Q4_K_S43.15 GiB46,331,287,6804.919RemySkye
Q4_145.95 GiB49,342,666,8805.239RemySkye
Q4_K_M48.06 GiB51,601,807,4885.478RemySkye
Q5_050.47 GiB54,197,049,4725.754RemySkye
Q5_K_S51.13 GiB54,901,692,5445.829RemySkye
Q5_K_M54.65 GiB58,681,006,2086.230RemySkye
Q5_155.00 GiB59,051,432,0646.269RemySkye
Q6_K62.63 GiB67,243,104,3847.139RemySkye
Q8_077.72 GiB83,453,368,4488.860RemySkye
BF16146.01 GiB156,772,423,80816.644RemySkye

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 39.48 GiB. The real file is 48.06 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
Attention heads
32
KV heads
2
Head dim
128
Hidden size
4096
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
512
Experts per token
1
use_sliding_window

Questions people ask

How much VRAM does NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 need?
Q4_K_M is exactly 51,601,807,488 bytes (48.06 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Is NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 a mixture-of-experts model?
Yes — 512 experts, 1 routed per token. Every expert must be resident, but only the routed ones are read per token, which is why its memory requirement and its speed behave very differently.
Which quantization of NVIDIA-Nemotron-Labs-3-Puzzle-75B-A9B-BF16 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.