canopylabs · audio tts

orpheus-3b-0.1-pretrained

canopylabs/orpheus-3b-0.1-pretrained

orpheus-3b-0.1-pretrained at Q4_K_M is exactly 2,363,760,640 bytes (2.20 GiB / 2.36 GB) — an effective 4.999 bits per weight, not the nominal 4. Its KV cache at 32K is 3.50 GiB.

From the file· summed from 1 file(s)From the file· KV from mirror (mirror:unsloth/orpheus-3b-0.1-pretrained)
Parameters
3.8B
Architecture
llama
28 layers
Context
131,072
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XXS1.40 GiB1,498,995,6163.170Mungert
IQ2_XS1.47 GiB1,579,113,3763.339Mungert
Q2_K1.49 GiB1,595,321,9203.374QuantFactory
IQ2_S1.50 GiB1,615,240,0963.416Mungert
IQ2_M1.55 GiB1,665,571,7443.522Mungert
Q2_K_S1.58 GiB1,694,817,1843.584Mungert
IQ3_XXS1.67 GiB1,796,119,4563.798Mungert
Q2_K_M1.70 GiB1,822,880,2243.855Mungert
Q3_K_S1.70 GiB1,823,200,4803.856QuantFactory
IQ3_XS1.71 GiB1,834,654,6243.880Mungert
Q3_K_M1.83 GiB1,967,510,7524.161QuantFactory
IQ3_M1.84 GiB1,975,032,7364.177Mungert
IQ3_S1.84 GiB1,975,032,7364.177Mungert
Q3_K_S1.88 GiB2,017,352,6084.266Mungert
Q2_K_L1.92 GiB2,056,406,9444.349Mungert
Q3_K_L1.95 GiB2,095,699,1684.432QuantFactory
Q4_01.99 GiB2,137,274,2724.520Mungert
Q3_K_M2.00 GiB2,145,415,6484.537Mungert
IQ4_XS2.01 GiB2,158,424,1284.564Mungert
IQ4_NL2.11 GiB2,261,570,7524.783Mungert
Q4_02.11 GiB2,261,573,6324.783QuantFactory
Q4_K_S2.12 GiB2,272,583,6804.806QuantFactory
Q4_K_M2.20 GiB2,363,760,6404.999QuantFactory
Q4_12.21 GiB2,373,700,0005.020Mungert
Q3_K_L2.22 GiB2,378,942,3685.031Mungert
Q4_K_S2.25 GiB2,414,572,0005.106Mungert
Q4_12.30 GiB2,467,866,8805.219QuantFactory
Q4_K_M2.30 GiB2,472,571,3605.229Mungert
Q5_02.43 GiB2,610,125,7285.520Mungert
Q5_02.49 GiB2,674,160,1285.655QuantFactory
Q5_K_S2.49 GiB2,674,160,1285.655QuantFactory
Q4_K_L2.52 GiB2,706,098,0805.723Mungert
Q5_K_M2.54 GiB2,726,801,9205.766QuantFactory
Q5_K_S2.58 GiB2,768,269,7925.854Mungert
Q5_K_M2.61 GiB2,798,350,8165.918Mungert
Q5_12.65 GiB2,846,551,4566.020Mungert
Q5_12.68 GiB2,880,453,3766.091QuantFactory
Q5_K_L2.82 GiB3,031,877,5366.412Mungert
Q6_K_M2.90 GiB3,112,530,4006.582Mungert
Q6_K2.90 GiB3,112,533,2806.582QuantFactory

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.44 GiB0.44 GiB28 / 0 / 0
8,1920.88 GiB0.88 GiB28 / 0 / 0
16,3841.75 GiB1.75 GiB28 / 0 / 0
32,7683.50 GiB3.50 GiB28 / 0 / 0
65,5367.00 GiB7.00 GiB28 / 0 / 0
131,07214.00 GiB14.00 GiB28 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 1.98 GiB. The real file is 2.20 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from mirror:unsloth/orpheus-3b-0.1-pretrained
Layers
28
Attention heads
24
KV heads
8
Head dim
128
Hidden size
3072
Vocab
156,939
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does orpheus-3b-0.1-pretrained need?
Q4_K_M is exactly 2,363,760,640 bytes (2.20 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is orpheus-3b-0.1-pretrained's KV cache?
3.50 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of orpheus-3b-0.1-pretrained should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.