deepseek-community · image

Janus-Pro-7B

deepseek-community/Janus-Pro-7B

Janus-Pro-7B at Q4_K_M is exactly 4,223,359,936 bytes (3.93 GiB / 4.22 GB) — an effective 4.561 bits per weight, not the nominal 4. Its KV cache at 32K is 15.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
7.4B
Architecture
llama
30 layers
Context
16,384
native (config.json)
License
mit

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S1.61 GiB1,733,033,1201.872mradermacher
I1-IQ1_M1.72 GiB1,848,564,8961.996mradermacher
I1-IQ2_XXS1.90 GiB2,041,117,8562.204mradermacher
I1-IQ2_XS2.06 GiB2,210,888,8642.388mradermacher
I1-IQ2_S2.23 GiB2,389,122,2082.580mradermacher
I1-Q2_K_S2.34 GiB2,510,511,2642.711mradermacher
I1-IQ2_M2.37 GiB2,543,164,5762.747mradermacher
Q2_K2.53 GiB2,718,424,0002.936mradermacher
I1-Q2_K2.53 GiB2,718,424,2242.936mradermacher
I1-IQ3_XXS2.57 GiB2,758,401,1842.979mradermacher
I1-IQ3_XS2.79 GiB2,993,609,8883.233mradermacher
Q3_K_S2.92 GiB3,138,018,2403.389mradermacher
I1-Q3_K_S2.92 GiB3,138,018,4643.389mradermacher
I1-IQ3_S2.92 GiB3,138,018,4643.389mradermacher
I1-IQ3_M3.06 GiB3,289,676,9603.553mradermacher
Q3_K_M3.22 GiB3,461,192,6403.738mradermacher
I1-Q3_K_M3.22 GiB3,461,192,8643.738mradermacher
Q3_K_L3.49 GiB3,746,274,2404.046mradermacher
I1-Q3_K_L3.49 GiB3,746,274,4644.046mradermacher
I1-IQ4_XS3.54 GiB3,797,228,7044.101mradermacher
IQ4_XS3.56 GiB3,818,363,8404.124mradermacher
I1-IQ4_NL3.73 GiB4,000,062,6244.320mradermacher
I1-Q4_03.73 GiB4,008,516,7684.329mradermacher
Q4_K_S3.75 GiB4,025,359,2964.347mradermacher
I1-Q4_K_S3.75 GiB4,025,359,5204.347mradermacher
Q4_K_M3.93 GiB4,223,359,9364.561mradermacher
I1-Q4_K_M3.93 GiB4,223,360,1604.561mradermacher
I1-Q4_14.10 GiB4,405,730,4644.758mradermacher
Q5_K_S4.48 GiB4,811,398,0805.196mradermacher
I1-Q5_K_S4.48 GiB4,811,398,3045.196mradermacher
Q5_K_M4.59 GiB4,926,430,1445.320mradermacher
I1-Q5_K_M4.59 GiB4,926,430,3685.320mradermacher
Q6_K5.28 GiB5,673,442,2406.127mradermacher
I1-Q6_K5.28 GiB5,673,442,4646.127mradermacher
Q8_06.84 GiB7,346,985,9207.934mradermacher
F1612.88 GiB13,825,219,52014.931mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.88 GiB1.88 GiB30 / 0 / 0
8,1923.75 GiB3.75 GiB30 / 0 / 0
16,3847.50 GiB7.50 GiB30 / 0 / 0
32,76815.00 GiB15.00 GiB30 / 0 / 0
65,53630.00 GiB30.00 GiB30 / 0 / 0
131,07260.00 GiB60.00 GiB30 / 0 / 0

Pipeline components

a diffusion model is a graph of parts, not one file
ComponentSizeShareCan live on the CPU?
denoiser13.80 GiB100%no, must be resident
Full pipeline13.80 GiBresident if nothing is offloaded

The parameter count published for a diffusion model describes the denoiser alone. Running it also requires its text encoder and VAE, and the text encoder is often nearly as large as the denoiser — which is why offloading it is the standard first move when you run out of memory.

We publish component sizes here, not throughput. Community-submitted image-generation rates do exist for many GPUs and we show them on the hardware pages, but they aggregate runs at different resolutions, step counts and settings, so they cannot be attributed to one model. Peak memory during sampling is unmeasured by any public source, and we do not estimate it.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 3.88 GiB. The real file is 3.93 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
30
Attention heads
32
KV heads
32
Head dim
128
Hidden size
4096
Vocab
102,400
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Janus-Pro-7B need?
Q4_K_M is exactly 4,223,359,936 bytes (3.93 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Janus-Pro-7B's KV cache?
15.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Janus-Pro-7B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.