Step-3.7-Flash
stepfun-ai/Step-3.7-FlashStep-3.7-Flash at Q4_K_M is exactly 121,621,146,016 bytes (113.27 GiB / 121.62 GB) — an effective 4.832 bits per weight, not the nominal 4. Its KV cache at 32K is 13.03 GiB, not the 45.00 GiB a flat formula predicts.
Shipped quantizations
| Quant | Size● | Exact bytes● | Effective bpw● | Tensors● | Publisher |
|---|---|---|---|---|---|
| IQ1_S | 40.03 GiB | 42,985,858,016 | 1.708 | — | bartowski |
| IQ1_M | 44.41 GiB | 47,681,234,912 | 1.894 | — | bartowski |
| IQ2_XXS2 shards | 51.22 GiB | 54,995,682,496 | 2.185 | — | bartowski |
| UD-IQ1_M3 shards | 52.85 GiB | 56,750,287,328 | 2.255 | — | unsloth |
| IQ2_XS2 shards | 56.76 GiB | 60,941,612,192 | 2.421 | — | bartowski |
| UD-IQ2_XXS3 shards | 57.47 GiB | 61,704,808,928 | 2.451 | — | unsloth |
| UD-IQ2_M3 shards | 57.58 GiB | 61,821,856,224 | 2.456 | — | unsloth |
| IQ2_S2 shards | 57.93 GiB | 62,196,901,024 | 2.471 | — | bartowski |
| IQ2_M2 shards | 63.68 GiB | 68,378,760,384 | 2.717 | — | bartowski |
| Q2_K2 shards | 66.34 GiB | 71,226,835,104 | 2.830 | — | bartowski |
| Q2_K_L2 shards | 66.82 GiB | 71,742,419,136 | 2.850 | — | bartowski |
| UD-IQ3_XXS3 shards | 68.54 GiB | 73,591,039,488 | 2.924 | — | unsloth |
| IQ3_XXS2 shards | 70.56 GiB | 75,758,680,960 | 3.010 | — | stepfun-ai |
| UD-IQ3_S3 shards | 74.54 GiB | 80,031,917,568 | 3.180 | — | unsloth |
| IQ3_XXS3 shards | 78.38 GiB | 84,162,180,384 | 3.344 | — | bartowski |
| Q3_K_S3 shards | 81.55 GiB | 87,566,150,976 | 3.479 | — | bartowski |
| UD-Q3_K_M3 shards | 83.13 GiB | 89,262,719,488 | 3.546 | — | unsloth |
| IQ3_XS3 shards | 85.44 GiB | 91,736,206,624 | 3.645 | — | bartowski |
| Q3_K_M3 shards | 85.51 GiB | 91,813,178,656 | 3.648 | — | bartowski |
| Q3_K_M3 shards | 87.36 GiB | 93,801,444,352 | 3.727 | — | stepfun-ai |
| UD-IQ4_XS3 shards | 88.79 GiB | 95,336,010,208 | 3.788 | — | unsloth |
| Q3_K_L3 shards | 88.94 GiB | 95,502,724,416 | 3.794 | — | bartowski |
| IQ3_M3 shards | 89.38 GiB | 95,976,156,480 | 3.813 | — | bartowski |
| UD-IQ4_NL3 shards | 90.63 GiB | 97,317,818,880 | 3.866 | — | unsloth |
| Q3_K_L3 shards | 95.46 GiB | 102,504,428,544 | 4.072 | — | stepfun-ai |
| IQ4_XS3 shards | 97.78 GiB | 104,993,562,624 | 4.171 | — | stepfun-ai |
| IQ4_XS3 shards | 99.94 GiB | 107,305,015,552 | 4.263 | — | bartowski |
| Q4_K_S3 shards | 103.84 GiB | 111,499,087,872 | 4.430 | — | stepfun-ai |
| IQ4_NL3 shards | 105.54 GiB | 113,325,345,088 | 4.502 | — | bartowski |
| Q4_03 shards | 105.79 GiB | 113,589,193,024 | 4.513 | — | bartowski |
| UD-Q4_K_S4 shards | 106.32 GiB | 114,163,192,448 | 4.536 | — | unsloth |
| Q4_K_S3 shards | 109.05 GiB | 117,089,601,824 | 4.652 | — | bartowski |
| Q4_K_M4 shards | 113.27 GiB | 121,621,146,016 | 4.832 | — | bartowski |
| Q4_K_L4 shards | 113.63 GiB | 122,012,989,856 | 4.847 | — | bartowski |
| UD-Q4_K_M4 shards | 113.71 GiB | 122,090,426,976 | 4.851 | — | unsloth |
| Q4_14 shards | 116.67 GiB | 125,278,316,992 | 4.977 | — | bartowski |
| Q5_K_S4 shards | 128.01 GiB | 137,450,531,232 | 5.461 | — | bartowski |
| UD-Q5_K_S4 shards | 128.59 GiB | 138,072,760,960 | 5.486 | — | unsloth |
| Q5_K_M4 shards | 132.28 GiB | 142,034,323,904 | 5.643 | — | bartowski |
| Q5_K_L4 shards | 132.58 GiB | 142,360,172,992 | 5.656 | — | bartowski |
KV cache by context
| Context | KV cache (f16)● | Flat formula | Overstated by | Full / windowed / recurrent |
|---|---|---|---|---|
| 4,096 | 2.53 GiB | 5.63 GiB | 2.22× | 12 / 33 / 0 |
| 8,192 | 4.03 GiB | 11.25 GiB | 2.79× | 12 / 33 / 0 |
| 16,384 | 7.03 GiB | 22.50 GiB | 3.20× | 12 / 33 / 0 |
| 32,768 | 13.03 GiB | 45.00 GiB | 3.45× | 12 / 33 / 0 |
| 65,536 | 25.03 GiB | 90.00 GiB | 3.60× | 12 / 33 / 0 |
| 131,072 | 49.03 GiB | 180.00 GiB | 3.67× | 12 / 33 / 0 |
33 of 45 layers cache only a 512-token window rather than the full context, on a period of . Figures assume the default configuration; --swa-full disables the saving entirely.
Compare with
Will it run on your card?
Why other calculators give a different number
A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 105.49 GiB. The real file is 113.27 GiB, because a quantization is a mixture and some tensors are always kept at higher precision. The larger discrepancy is the cache: a flat formula gives 45.00 GiB at 32K context where the real figure is 13.03 GiB, because most of this model's layers cache a fixed window rather than the whole context.
Architecture
Questions people ask
- How much VRAM does Step-3.7-Flash need?
- Q4_K_M is exactly 121,621,146,016 bytes (113.27 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
- How large is Step-3.7-Flash's KV cache?
- 13.03 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
- Which quantization of Step-3.7-Flash should I use?
- Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.