Wan-AI · text

Wan2.1-T2V-1.3B

Wan-AI/Wan2.1-T2V-1.3B

Wan2.1-T2V-1.3B at Q4_K_M is exactly 982,716,640 bytes (0.92 GiB / 0.98 GB) — an effective 5.540 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
1.4B
Architecture
qwen35
null layers
Context
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q3_K_S0.61 GiB654,903,5203.692samuelchristlie
Q3_K_M0.68 GiB729,129,1844.111825samuelchristlie
Q4_00.81 GiB865,581,2804.880825samuelchristlie
Q4_K_S0.83 GiB892,565,7285.032samuelchristlie
Q4_10.86 GiB926,775,5205.225samuelchristlie
Q4_K_M0.92 GiB982,716,6405.540825samuelchristlie
Q5_K_S0.94 GiB1,013,774,5605.715samuelchristlie
Q3_K_S0.95 GiB1,024,104,1285.774calcuis
Q5_00.97 GiB1,039,579,3605.861samuelchristlie
Q5_K_M1.01 GiB1,087,410,4006.131825samuelchristlie
Q5_11.03 GiB1,100,773,6006.206samuelchristlie
Q6_K1.12 GiB1,198,647,5206.758825samuelchristlie
Q3_K_L1.15 GiB1,238,514,3686.982calcuis
Q4_K_S1.29 GiB1,385,021,1207.809calcuis
Q2_K2 shards1.31 GiB1,408,997,2807.944calcuis
Q8_01.43 GiB1,535,768,8008.658825samuelchristlie
Q5_K_S1.46 GiB1,572,142,7848.863calcuis
Q3_K_M2 shards1.78 GiB1,915,591,58410.800calcuis
Q4_K_M2.54 GiB2,724,723,26415.361427DeepBeepMeep
Q5_K_M2 shards2.63 GiB2,821,321,63215.906calcuis
F162.65 GiB2,840,754,40016.016825samuelchristlie
Q4_K_M4.69 GiB5,038,065,21628.404DeepBeepMeep
Q4_14 shards4.94 GiB5,299,287,52029.876calcuis
Q4_K_M3 shards5.78 GiB6,210,013,952calcuis
Q5_14 shards5.84 GiB6,270,432,736calcuis
Q6_K4 shards6.35 GiB6,816,701,920calcuis
Q4_06 shards6.50 GiB6,980,190,944calcuis
BF162 shards6.76 GiB7,254,226,848calcuis
F163 shards6.99 GiB7,508,102,752calcuis
Q5_06 shards7.76 GiB8,334,721,760calcuis
Q8_06 shards11.37 GiB12,204,778,208calcuis

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 0.74 GiB. The real file is 0.92 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
Attention heads
KV heads
Head dim
Hidden size
Vocab
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Wan2.1-T2V-1.3B need?
Q4_K_M is exactly 982,716,640 bytes (0.92 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Wan2.1-T2V-1.3B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.