BlinkDL · text

rwkv7-g1

BlinkDL/rwkv7-g1

rwkv7-g1 at I1-Q3_K_M is exactly 1,760,123,040 bytes (1.64 GiB / 1.76 GB) — an effective 1.061 bits per weight, not the nominal 1.

From the file· summed from 1 file(s)
Parameters
13.3B
Architecture
rwkv7
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-Q3_K_M1.64 GiB1,760,123,0401.061RemySkye
I1-IQ4_XS1.67 GiB1,791,908,0321.080RemySkye
I1-IQ4_NL1.75 GiB1,875,794,1121.131RemySkye
I1-Q4_K_S1.75 GiB1,875,794,1121.131RemySkye
I1-Q4_01.75 GiB1,875,794,1121.131RemySkye
I1-Q3_K_L1.78 GiB1,909,217,4401.151RemySkye
I1-Q4_11.90 GiB2,043,566,2721.232RemySkye
I1-Q4_K_M1.91 GiB2,054,215,8401.238RemySkye
I1-Q5_K_S2.06 GiB2,211,338,4321.333RemySkye
I1-Q5_02.06 GiB2,211,338,4321.333RemySkye
I1-Q5_K_M2.15 GiB2,303,252,6401.389RemySkye
I1-Q5_12.22 GiB2,379,110,5921.434RemySkye
I1-Q6_K2.39 GiB2,567,854,2721.548RemySkye
I1-IQ1_S3.52 GiB3,780,198,6242.279RemySkye
I1-IQ1_M3.79 GiB4,068,032,7362.452RemySkye
I1-IQ2_XXS4.24 GiB4,547,756,2562.741RemySkye
I1-IQ2_XS4.59 GiB4,931,535,0722.973RemySkye
I1-IQ2_S4.62 GiB4,958,798,0482.989RemySkye
I1-IQ2_M4.98 GiB5,342,576,8643.221RemySkye
I1-Q2_K5.07 GiB5,446,910,1763.284RemySkye
I1-IQ3_XXS5.69 GiB6,110,134,4963.683RemySkye
I1-IQ3_S6.26 GiB6,721,454,3044.052RemySkye
I1-Q3_K_S6.26 GiB6,721,454,3044.052RemySkye
I1-Q3_K_M7.14 GiB7,671,201,9524.624RemySkye
I1-IQ4_XS7.45 GiB7,995,998,4324.820RemySkye
I1-Q4_K_S7.81 GiB8,388,165,8565.057RemySkye
I1-IQ4_NL7.81 GiB8,388,165,8565.057RemySkye
I1-Q4_07.81 GiB8,388,165,8565.057RemySkye
I1-Q3_K_L7.83 GiB8,409,399,4565.069RemySkye
I1-Q4_K_M8.48 GiB9,106,178,2085.489RemySkye
I1-Q4_18.54 GiB9,172,500,7045.529RemySkye
I1-Q5_K_S9.27 GiB9,956,835,5526.002RemySkye
I1-Q5_09.27 GiB9,956,835,5526.002RemySkye
I1-Q5_K_M9.62 GiB10,326,720,6726.225RemySkye
I1-Q5_110.00 GiB10,741,170,4006.475RemySkye
I1-Q6_K10.83 GiB11,623,547,1047.007RemySkye

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts I1-Q3_K_M at roughly 6.95 GiB. The real file is 1.64 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does rwkv7-g1 need?
I1-Q3_K_M is exactly 1,760,123,040 bytes (1.64 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of rwkv7-g1 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.