empero-ai · text

openNemo-9B-abliterated

empero-ai/openNemo-9B-abliterated

openNemo-9B-abliterated at Q4_K_M is exactly 6,525,626,432 bytes (6.08 GiB / 6.53 GB) — an effective 5.873 bits per weight, not the nominal 4. Its KV cache at 32K is 7.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
8.9B
Architecture
nemotron_h
56 layers
Context
131,072
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K4.66 GiB5,006,189,6324.506empero-ai
Q3_K_S4.78 GiB5,131,988,0324.619empero-ai
Q4_04.94 GiB5,308,679,2324.778empero-ai
IQ4_XS4.99 GiB5,359,303,2324.824empero-ai
Q3_K_M5.01 GiB5,379,732,0324.842empero-ai
Q3_K_L5.11 GiB5,488,363,0724.940empero-ai
Q4_K_S5.79 GiB6,211,614,2725.591empero-ai
Q5_05.91 GiB6,346,032,1925.712empero-ai
Q4_K_M6.08 GiB6,525,626,4325.873empero-ai
Q5_K_S6.32 GiB6,781,559,8726.104empero-ai
Q5_K_M6.58 GiB7,069,803,0726.363empero-ai
Q6_K8.51 GiB9,135,889,4728.223empero-ai
Q8_08.81 GiB9,458,091,0728.513empero-ai
BF1616.57 GiB17,788,740,67216.011empero-ai

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.88 GiB0.88 GiB56 / 0 / 0
8,1921.75 GiB1.75 GiB56 / 0 / 0
16,3843.50 GiB3.50 GiB56 / 0 / 0
32,7687.00 GiB7.00 GiB56 / 0 / 0
65,53614.00 GiB14.00 GiB56 / 0 / 0
131,07228.00 GiB28.00 GiB56 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.66 GiB. The real file is 6.08 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
56
Attention heads
40
KV heads
8
Head dim
128
Hidden size
4480
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does openNemo-9B-abliterated need?
Q4_K_M is exactly 6,525,626,432 bytes (6.08 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is openNemo-9B-abliterated's KV cache?
7.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of openNemo-9B-abliterated should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.