tencent · text

Hunyuan-7B-Instruct

tencent/Hunyuan-7B-Instruct

Hunyuan-7B-Instruct at Q4_K_M is exactly 4,621,992,928 bytes (4.30 GiB / 4.62 GB) — an effective 4.927 bits per weight, not the nominal 4. Its KV cache at 32K is 4.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
7.5B
Architecture
hunyuan-dense
32 layers
Context
32,768
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XXS2.07 GiB2,223,644,9282.370gabriellarson
IQ2_XS2.26 GiB2,430,214,4002.591gabriellarson
IQ2_S2.36 GiB2,529,566,9762.697gabriellarson
IQ2_M2.53 GiB2,719,359,2322.899gabriellarson
IQ2_M2.53 GiB2,719,359,2642.899bartowski
Q2_K_S2.62 GiB2,813,199,3282.999gabriellarson
Q2_K2.80 GiB3,003,515,8723.202gabriellarson
Q2_K2.80 GiB3,003,515,9043.202bartowski
IQ3_XXS2.84 GiB3,045,990,6563.247gabriellarson
IQ3_XXS2.84 GiB3,045,990,6883.247bartowski
Q2_K_L2.92 GiB3,130,657,5683.337bartowski
IQ3_XS3.06 GiB3,289,777,1203.507gabriellarson
IQ3_XS3.06 GiB3,289,777,1523.507bartowski
Q3_K_S3.20 GiB3,435,529,1843.662gabriellarson
Q3_K_S3.20 GiB3,435,529,2163.662bartowski
IQ3_S3.22 GiB3,453,354,9763.681gabriellarson
IQ3_M3.31 GiB3,555,853,2803.791gabriellarson
IQ3_M3.31 GiB3,555,853,3123.791bartowski
Q3_K_M3.53 GiB3,789,947,8724.040gabriellarson
Q3_K_M3.53 GiB3,789,947,9044.040bartowski
Q3_K_L3.81 GiB4,092,986,3364.363gabriellarson
Q3_K_L3.81 GiB4,092,986,3684.363bartowski
IQ4_XS3.88 GiB4,165,338,0804.440gabriellarson
IQ4_XS3.88 GiB4,165,338,1124.440bartowski
Q4_04.08 GiB4,377,150,4324.666gabriellarson
Q4_04.08 GiB4,377,150,4644.666bartowski
IQ4_NL4.08 GiB4,379,247,5844.668gabriellarson
IQ4_NL4.08 GiB4,379,247,6164.668bartowski
Q4_K_S4.09 GiB4,393,927,6484.684gabriellarson
Q4_K_S4.09 GiB4,393,927,6804.684bartowski
Q4_K_M4.30 GiB4,621,992,9284.927gabriellarson
Q4_K_M4.30 GiB4,621,992,9604.927bartowski
Q4_K_L4.42 GiB4,749,134,6245.063bartowski
Q4_14.47 GiB4,798,678,0165.115bartowski
Q5_K_S4.88 GiB5,234,885,6005.580gabriellarson
Q5_K_S4.88 GiB5,234,885,6325.580bartowski
Q5_04.89 GiB5,249,565,6645.596gabriellarson
Q5_K_M5.00 GiB5,368,579,0405.723gabriellarson
Q5_K_M5.00 GiB5,368,579,0725.723bartowski
Q5_K_L5.12 GiB5,495,720,7365.859bartowski

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.50 GiB0.50 GiB32 / 0 / 0
8,1921.00 GiB1.00 GiB32 / 0 / 0
16,3842.00 GiB2.00 GiB32 / 0 / 0
32,7684.00 GiB4.00 GiB32 / 0 / 0
65,5368.00 GiB8.00 GiB32 / 0 / 0
131,07216.00 GiB16.00 GiB32 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 3.93 GiB. The real file is 4.30 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
32
Attention heads
32
KV heads
8
Head dim
128
Hidden size
4096
Vocab
128,167
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Hunyuan-7B-Instruct need?
Q4_K_M is exactly 4,621,992,928 bytes (4.30 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Hunyuan-7B-Instruct's KV cache?
4.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Hunyuan-7B-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.