arcee-ai · text

Homunculus

arcee-ai/Homunculus

Homunculus at Q4_K_M is exactly 7,621,176,448 bytes (7.10 GiB / 7.62 GB) — an effective 4.894 bits per weight, not the nominal 4. Its KV cache at 32K is 5.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
12.5B
Architecture
llama
40 layers
Context
131,072
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XS3.74 GiB4,020,306,0482.582bartowski
IQ2_S3.96 GiB4,254,418,0482.732bartowski
IQ2_M4.24 GiB4,550,968,4482.922bartowski
Q2_K4.57 GiB4,910,290,0483.153bartowski
IQ3_XXS4.71 GiB5,061,330,0483.250bartowski
IQ3_XS5.06 GiB5,436,446,8483.491bartowski
Q3_K_S5.28 GiB5,664,184,4483.637bartowski
Q2_K_L5.28 GiB5,668,690,0483.640bartowski
IQ3_M5.45 GiB5,852,190,8483.758bartowski
Q3_K_M5.79 GiB6,213,048,4483.990bartowski
Q3_K_L6.23 GiB6,691,461,2484.297bartowski
IQ4_XS6.41 GiB6,883,384,4484.420bartowski
Q4_06.74 GiB7,238,610,0484.648bartowski
IQ4_NL6.74 GiB7,241,886,8484.650bartowski
Q4_K_S6.77 GiB7,264,169,0884.664bartowski
Q4_K_M7.10 GiB7,621,176,4484.894bartowski
Q4_17.40 GiB7,945,784,4485.102bartowski
Q4_K_L7.63 GiB8,197,560,4485.264bartowski
Q5_K_S8.08 GiB8,675,896,4485.571bartowski
Q5_K_M8.27 GiB8,884,792,4485.705bartowski
Q5_K_L8.72 GiB9,364,101,2486.013bartowski
Q6_K9.52 GiB10,227,384,4486.567bartowski
Q6_K_L9.88 GiB10,603,550,8486.809bartowski
Q8_012.34 GiB13,244,651,6488.505bartowski
BF1623.21 GiB24,924,395,36016.004bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.63 GiB0.63 GiB40 / 0 / 0
8,1921.25 GiB1.25 GiB40 / 0 / 0
16,3842.50 GiB2.50 GiB40 / 0 / 0
32,7685.00 GiB5.00 GiB40 / 0 / 0
65,53610.00 GiB10.00 GiB40 / 0 / 0
131,07220.00 GiB20.00 GiB40 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 6.53 GiB. The real file is 7.10 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
40
Attention heads
32
KV heads
8
Head dim
128
Hidden size
5120
Vocab
151,680
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Homunculus need?
Q4_K_M is exactly 7,621,176,448 bytes (7.10 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Homunculus's KV cache?
5.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Homunculus should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.