nvidia · embedding

Nemotron-3-Embed-8B-BF16

nvidia/Nemotron-3-Embed-8B-BF16

Nemotron-3-Embed-8B-BF16 at Q4_K_M is exactly 4,896,389,984 bytes (4.56 GiB / 4.90 GB) — an effective 4.926 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
8.0B
Architecture
mistral3
34 layers
Context
262,144
native (config.json)
License
openmdw-1.1

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S1.81 GiB1,945,664,3521.957Aqua00
IQ1_M1.95 GiB2,097,249,1202.110Aqua00
IQ2_XXS2.19 GiB2,349,890,4002.364Aqua00
Q2_02.36 GiB2,535,029,6002.550Aqua00
IQ2_XS2.39 GiB2,569,829,2162.585Aqua00
IQ2_S2.49 GiB2,673,900,3842.690Aqua00
IQ2_M2.68 GiB2,876,013,4082.893Aqua00
Q2_K_S2.77 GiB2,971,106,1442.989Aqua00
Q2_K2.96 GiB3,176,758,1123.196Aqua00
IQ3_XXS3.00 GiB3,224,664,9283.244Aqua00
IQ3_XS3.24 GiB3,483,663,2003.504Aqua00
Q3_K_S3.39 GiB3,635,772,2563.657Aqua00
IQ3_S3.40 GiB3,654,712,1603.676Aqua00
IQ3_M3.50 GiB3,761,666,9123.784Aqua00
Q3_K_M3.74 GiB4,011,359,0724.035Aqua00
Q3_K_M3.74 GiB4,011,359,1274.035Abiray
Q3_K_L4.04 GiB4,334,320,4804.360Aqua00
IQ4_XS4.11 GiB4,411,194,2084.437Aqua00
Q4_04.32 GiB4,635,327,3284.663Aqua00
IQ4_NL4.32 GiB4,638,473,0564.666Aqua00
Q4_K_S4.33 GiB4,652,104,5444.680Aqua00
Q4_K_S4.33 GiB4,652,104,5994.680Abiray
Q4_K_M4.56 GiB4,896,389,9844.926Aqua00
Q4_K_M4.56 GiB4,896,390,0394.926Abiray
Q4_14.73 GiB5,084,117,8565.114Aqua00
Q5_K_S5.17 GiB5,547,588,4485.581Aqua00
Q5_K_S5.17 GiB5,547,588,5035.581Abiray
Q5_05.18 GiB5,562,268,5125.595Aqua00
Q5_K_M5.30 GiB5,689,637,7285.723Aqua00
Q5_K_M5.30 GiB5,689,637,7835.723Abiray
Q5_15.60 GiB6,011,059,0406.047Aqua00
Q6_K6.08 GiB6,532,463,4566.571Aqua00
Q6_K6.08 GiB6,532,463,5096.571Abiray
Q8_07.88 GiB8,458,435,4248.509Aqua00
Q8_07.88 GiB8,458,435,4778.509Abiray
BF1614.82 GiB15,913,810,49616.009Aqua00

No KV cache

architectural, not a gap in our data

This architecture allocates no KV cache. Encoder and embedding models process their input in one pass rather than generating token by token, so there is nothing to carry forward between steps and memory does not grow with context. Its footprint is the weights plus a working buffer, and that is the whole story.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.17 GiB. The real file is 4.56 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
34
Attention heads
32
KV heads
8
Head dim
128
Hidden size
4096
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Nemotron-3-Embed-8B-BF16 need?
Q4_K_M is exactly 4,896,389,984 bytes (4.56 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Nemotron-3-Embed-8B-BF16 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.