arcee-ai · text · mixture of experts

Trinity-Large-Preview

arcee-ai/Trinity-Large-Preview

Trinity-Large-Preview at Q4_K_M is exactly 239,855,587,136 bytes (223.38 GiB / 239.86 GB) — an effective 4.814 bits per weight, not the nominal 4. Its KV cache at 32K is 2.67 GiB, not the 7.50 GiB a flat formula predicts.

From the file· summed from 5 file(s)From the file· KV per layer
Parameters
399B
total, not active
Architecture
afmoe
60 layers
Context
262,144
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S3 shards76.06 GiB81,666,719,1041.639arcee-ai
IQ1_S3 shards76.06 GiB81,666,719,1041.639bartowski
IQ1_M3 shards79.29 GiB85,135,768,9281.708bartowski
IQ1_M3 shards79.29 GiB85,135,768,9281.708arcee-ai
IQ2_XXS3 shards88.58 GiB95,107,628,4161.909bartowski
IQ2_XXS3 shards88.58 GiB95,107,628,4161.909arcee-ai
IQ2_XS3 shards102.27 GiB109,811,940,7362.204arcee-ai
IQ2_XS3 shards102.27 GiB109,811,940,7362.204bartowski
IQ2_S3 shards102.49 GiB110,050,463,1042.208arcee-ai
IQ2_S3 shards102.49 GiB110,050,463,1042.208bartowski
IQ2_M4 shards116.50 GiB125,094,514,1442.510arcee-ai
IQ2_M4 shards116.50 GiB125,094,514,1442.510bartowski
Q2_K4 shards129.44 GiB138,989,784,5442.789bartowski
Q2_K4 shards129.44 GiB138,989,784,5442.789arcee-ai
Q2_K_L4 shards130.00 GiB139,590,360,5442.801arcee-ai
Q2_K_L4 shards130.00 GiB139,590,360,5442.801bartowski
Q2_K3 shards135.04 GiB144,993,750,5922.910unsloth
Q2_K_L3 shards135.17 GiB145,137,888,8322.913unsloth
UD-IQ2_XXS4 shards142.94 GiB153,484,316,3523.080unsloth
UD-IQ2_M4 shards143.68 GiB154,274,975,4243.096unsloth
IQ3_XXS4 shards146.20 GiB156,981,431,7763.150bartowski
IQ3_XXS4 shards146.20 GiB156,981,431,7763.150arcee-ai
IQ3_XS5 shards151.27 GiB162,428,149,3763.260arcee-ai
IQ3_XS5 shards151.27 GiB162,428,149,3763.260bartowski
Q3_K_S4 shards160.21 GiB172,025,028,2883.452unsloth
Q3_K_S5 shards160.91 GiB172,772,187,7443.467arcee-ai
Q3_K_S5 shards160.91 GiB172,772,187,7443.467bartowski
IQ3_M5 shards168.61 GiB181,045,943,8723.633arcee-ai
IQ3_M5 shards168.61 GiB181,045,943,8723.633bartowski
Q3_K_M5 shards168.74 GiB181,184,749,1203.636arcee-ai
Q3_K_M5 shards168.74 GiB181,184,749,1203.636bartowski
UD-IQ3_XXS4 shards170.01 GiB182,546,886,3363.663unsloth
Q3_K_L5 shards175.82 GiB188,781,485,6643.789arcee-ai
Q3_K_L5 shards175.82 GiB188,781,485,6643.789bartowski
Q3_K_M4 shards176.43 GiB189,440,417,5043.802unsloth
IQ4_XS5 shards197.87 GiB212,465,413,9524.264unsloth
IQ4_XS6 shards198.19 GiB212,807,904,9604.271bartowski
IQ4_XS6 shards198.19 GiB212,807,904,9604.271arcee-ai
IQ4_NL5 shards209.39 GiB224,829,304,6084.512unsloth
Q4_05 shards209.52 GiB224,969,019,1684.515unsloth

KV cache by context

computed per layer — this model uses sliding-window attention
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.94 GiB0.94 GiB15 / 45 / 0
8,1921.26 GiB1.88 GiB1.49×15 / 45 / 0
16,3841.73 GiB3.75 GiB2.17×15 / 45 / 0
32,7682.67 GiB7.50 GiB2.81×15 / 45 / 0
65,5364.54 GiB15.00 GiB3.30×15 / 45 / 0
131,0728.29 GiB30.00 GiB3.62×15 / 45 / 0

45 of 60 layers cache only a 4,096-token window rather than the full context, on a period of 4. Figures assume the default configuration; --swa-full disables the saving entirely.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 208.83 GiB. The real file is 223.38 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 7.50 GiB at 32K context where the real figure is 2.67 GiB, because most of this model's layers cache a fixed window rather than the whole context.

Architecture

from config.json
Layers
60
Attention heads
48
KV heads
8
Head dim
128
Hidden size
3072
Vocab
200,192
Sliding window
4096
SWA period
4
MLA
no
Experts
256
Experts per token
4
use_sliding_window

Questions people ask

How much VRAM does Trinity-Large-Preview need?
Q4_K_M is exactly 239,855,587,136 bytes (223.38 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Trinity-Large-Preview's KV cache?
2.67 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Is Trinity-Large-Preview a mixture-of-experts model?
Yes — 256 experts, 4 routed per token. Every expert must be resident, but only the routed ones are read per token, which is why its memory requirement and its speed behave very differently.
Which quantization of Trinity-Large-Preview should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.