OpenGVLab · vision language

InternVL3_5-30B-A3B

OpenGVLab/InternVL3_5-30B-A3B

InternVL3_5-30B-A3B at Q4_K_M is exactly 18,632,180,352 bytes (17.35 GiB / 18.63 GB) — an effective 4.832 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
30.8B
Architecture
qwen3moe
null layers
Context
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_XXS7.05 GiB7,569,021,8561.963bartowski
IQ2_XS8.07 GiB8,662,162,3362.246bartowski
IQ2_S8.14 GiB8,744,096,6722.268bartowski
IQ2_M9.19 GiB9,870,267,2962.560bartowski
Q2_K10.16 GiB10,908,644,2562.829bartowski
Q2_K_L10.44 GiB11,212,516,2562.908bartowski
IQ3_XXS11.38 GiB12,216,980,3843.168bartowski
IQ3_XS11.86 GiB12,736,457,6323.303bartowski
Q3_K_S12.51 GiB13,428,124,5763.482bartowski
IQ3_M13.11 GiB14,076,537,7603.651bartowski
Q3_K_M13.11 GiB14,076,799,9043.651bartowski
Q3_K_L13.58 GiB14,582,999,6803.782lmstudio-community
Q3_K_L13.58 GiB14,582,999,9683.782bartowski
IQ4_XS15.33 GiB16,457,999,2644.268bartowski
IQ4_NL16.19 GiB17,386,275,7444.509bartowski
Q4_016.42 GiB17,631,642,5284.572bartowski
Q4_K_S16.75 GiB17,984,488,3524.664bartowski
Q4_K_M17.35 GiB18,632,180,3524.832lmstudio-community
Q4_K_M17.35 GiB18,632,180,6404.832bartowski
Q4_K_L17.57 GiB18,863,123,3604.892bartowski
Q4_117.89 GiB19,214,517,1524.983bartowski
Q5_K_S19.65 GiB21,099,381,6645.472bartowski
Q5_K_M20.25 GiB21,744,452,5125.639bartowski
Q5_K_L20.43 GiB21,936,499,6165.689bartowski
Q6_K23.38 GiB25,104,718,4646.510lmstudio-community
Q6_K23.38 GiB25,104,718,7526.510bartowski
Q6_K_L23.52 GiB25,255,439,2646.550bartowski
Q8_030.25 GiB32,483,928,7048.424lmstudio-community
Q8_030.25 GiB32,483,928,9928.424bartowski
BF162 shards56.90 GiB61,095,799,64815.844bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 16.16 GiB. The real file is 17.35 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
Attention heads
KV heads
Head dim
Hidden size
Vocab
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does InternVL3_5-30B-A3B need?
Q4_K_M is exactly 18,632,180,352 bytes (17.35 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of InternVL3_5-30B-A3B should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.