ai9stars · text

G9v3-3B

ai9stars/G9v3-3B

G9v3-3B at Q4_K_M is exactly 1,843,804,448 bytes (1.72 GiB / 1.84 GB) — an effective 4.936 bits per weight, not the nominal 4. Its KV cache at 32K is 1.63 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
3.0B
Architecture
llama
52 layers
Context
131,072
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M1.15 GiB1,233,615,4243.302bartowski
Q2_K1.18 GiB1,266,866,7523.391bartowski
IQ3_XXS1.22 GiB1,306,860,0963.498bartowski
IQ3_XS1.30 GiB1,394,768,4483.733bartowski
Q3_K_S1.34 GiB1,437,153,8563.847bartowski
IQ3_M1.38 GiB1,483,176,5123.970bartowski
Q2_K_L1.42 GiB1,527,986,7524.090bartowski
Q3_K_M1.46 GiB1,563,064,8964.184bartowski
Q3_K_L1.53 GiB1,641,839,1684.395bartowski
IQ4_XS1.59 GiB1,710,316,0964.578bartowski
Q4_01.64 GiB1,755,945,2484.700WhiskyAKM
Q4_01.64 GiB1,755,945,2484.700WhiskyAKM
Q4_K_S1.64 GiB1,765,644,5764.726WhiskyAKM
Q4_K_S1.64 GiB1,765,644,5764.726WhiskyAKM
Q4_01.67 GiB1,788,779,0724.788bartowski
IQ4_NL1.67 GiB1,791,089,2164.794bartowski
Q4_K_S1.67 GiB1,795,201,6004.805bartowski
Q4_K_M1.72 GiB1,843,804,4484.936WhiskyAKM
Q4_K_M1.72 GiB1,843,804,4484.936WhiskyAKM
Q4_K_M1.77 GiB1,902,590,5285.093bartowski
Q4_11.81 GiB1,947,310,6565.213bartowski
Q5_K_S1.95 GiB2,096,077,0885.611WhiskyAKM
Q5_K_S1.95 GiB2,096,077,0885.611WhiskyAKM
Q4_K_L1.96 GiB2,101,041,7285.624bartowski
Q5_K_S1.97 GiB2,110,560,8325.649bartowski
Q5_K_M1.99 GiB2,141,337,8885.732WhiskyAKM
Q5_K_M1.99 GiB2,141,337,8885.732WhiskyAKM
Q5_K_M2.04 GiB2,188,409,4085.858bartowski
Q5_K_L2.19 GiB2,353,437,2486.300bartowski
Q6_K2.29 GiB2,457,467,1686.578WhiskyAKM
Q6_K2.29 GiB2,457,467,1686.578WhiskyAKM
Q6_K2.37 GiB2,546,604,6086.817bartowski
Q6_K_L2.49 GiB2,676,120,1287.163bartowski
Q8_02.96 GiB3,181,221,2808.515Szeweq
Q8_02.96 GiB3,181,230,3688.515WhiskyAKM
Q8_02.96 GiB3,181,230,3688.515WhiskyAKM
Q8_02.96 GiB3,181,230,6568.515bartowski
BF165.57 GiB5,982,894,36816.015bartowski
BF165.57 GiB5,982,894,36816.015WhiskyAKM
BF165.57 GiB5,982,894,36816.015WhiskyAKM

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.20 GiB0.20 GiB52 / 0 / 0
8,1920.41 GiB0.41 GiB52 / 0 / 0
16,3840.81 GiB0.81 GiB52 / 0 / 0
32,7681.63 GiB1.63 GiB52 / 0 / 0
65,5363.25 GiB3.25 GiB52 / 0 / 0
131,0726.50 GiB6.50 GiB52 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 1.57 GiB. The real file is 1.72 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
52
Attention heads
16
KV heads
2
Head dim
128
Hidden size
2048
Vocab
130,560
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does G9v3-3B need?
Q4_K_M is exactly 1,843,804,448 bytes (1.72 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is G9v3-3B's KV cache?
1.63 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of G9v3-3B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.