unsloth · text

Phi-3.5-mini-instruct

unsloth/Phi-3.5-mini-instruct

Phi-3.5-mini-instruct at IQ1_S is exactly 881,723,232 bytes (0.82 GiB / 0.88 GB) — an effective 1.846 bits per weight, not the nominal 1. Its KV cache at 32K is 12.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
3.8B
Architecture
llama
32 layers
Context
131,072
native (config.json)
License
mit

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S0.82 GiB881,723,2321.846legraphista
IQ1_M0.88 GiB950,142,8161.989legraphista
IQ2_XXS0.99 GiB1,064,175,4562.228legraphista
IQ2_XS1.08 GiB1,164,838,7522.439legraphista
IQ2_S1.17 GiB1,258,204,5122.634legraphista
Q2_K_S1.24 GiB1,327,342,9442.779legraphista
IQ2_M1.26 GiB1,349,430,6242.825legraphista
Q2_K1.35 GiB1,446,880,6083.029legraphista
IQ3_XXS1.37 GiB1,475,259,7443.089legraphista
IQ3_XS1.49 GiB1,596,869,4723.343legraphista
Q3_K_S1.57 GiB1,681,804,1283.521legraphista
IQ3_S1.57 GiB1,681,804,1283.521legraphista
IQ3_M1.65 GiB1,775,389,5363.717legraphista
Q3_K1.75 GiB1,877,625,6963.931legraphista
Q3_K_L1.90 GiB2,045,135,7124.282legraphista
IQ4_XS1.92 GiB2,059,858,2724.313legraphista
IQ4_NL2.03 GiB2,176,182,6244.556legraphista
Q4_K_S2.04 GiB2,193,484,1284.592legraphista
Q4_K2.16 GiB2,318,920,0324.855legraphista
Q5_K_S2.46 GiB2,641,479,7765.530legraphista
Q5_K2.53 GiB2,715,011,1685.684legraphista
Q6_K2.92 GiB3,135,858,2726.565legraphista
Q8_03.78 GiB4,061,227,6168.503legraphista
BF167.12 GiB7,643,302,49616.002legraphista

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.50 GiB1.50 GiB32 / 0 / 0
8,1923.00 GiB3.00 GiB32 / 0 / 0
16,3846.00 GiB6.00 GiB32 / 0 / 0
32,76812.00 GiB12.00 GiB32 / 0 / 0
65,53624.00 GiB24.00 GiB32 / 0 / 0
131,07248.00 GiB48.00 GiB32 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts IQ1_S at roughly 2.00 GiB. The real file is 0.82 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
32
Attention heads
32
KV heads
32
Head dim
96
Hidden size
3072
Vocab
32,064
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Phi-3.5-mini-instruct need?
IQ1_S is exactly 881,723,232 bytes (0.82 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Phi-3.5-mini-instruct's KV cache?
12.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Phi-3.5-mini-instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.