openbmb · text

MiniCPM-V-2_6

openbmb/MiniCPM-V-2_6

MiniCPM-V-2_6 at Q4_K_M is exactly 4,681,089,344 bytes (4.36 GiB / 4.68 GB) — an effective 4.624 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
8.1B
Architecture
qwen2
Context
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M2.59 GiB2,778,737,4402.745bartowski
Q2_K2.81 GiB3,014,290,2402.977openbmb
Q2_K2.81 GiB3,014,290,6882.977bartowski
IQ3_XS3.11 GiB3,344,461,1523.304openbmb
IQ3_XS3.11 GiB3,344,461,6003.304bartowski
Q3_K_S3.25 GiB3,490,573,6643.448openbmb
Q3_K_S3.25 GiB3,490,574,1123.448bartowski
IQ3_S3.26 GiB3,497,397,6003.455openbmb
Q2_K_L3.30 GiB3,545,121,6643.502bartowski
IQ3_M3.33 GiB3,572,217,1843.529openbmb
IQ3_M3.33 GiB3,572,217,6323.529bartowski
Q3_K_M3.55 GiB3,806,596,4483.760openbmb
Q3_K3.55 GiB3,806,596,4483.760openbmb
Q3_K_M3.55 GiB3,806,596,8963.760bartowski
Q3_K_L3.81 GiB4,086,664,5444.037openbmb
Q3_K_L3.81 GiB4,086,664,9924.037bartowski
IQ4_XS3.93 GiB4,216,533,2804.165bartowski
IQ4_XS3.96 GiB4,248,358,7524.196openbmb
Q4_04.13 GiB4,429,406,5284.375openbmb
Q4_K_S4.15 GiB4,455,784,7684.401openbmb
Q4_K_S4.15 GiB4,455,785,2164.401bartowski
IQ4_NL4.15 GiB4,461,289,7924.407openbmb
Q4_K_M4.36 GiB4,681,089,3444.624openbmb
Q4_K4.36 GiB4,681,089,3444.624openbmb
Q4_K_M4.36 GiB4,681,089,5364.624lmstudio-community
Q4_K_M4.36 GiB4,681,089,7924.624bartowski
Q4_14.54 GiB4,871,210,2404.812openbmb
Q4_K_L4.74 GiB5,084,521,3445.022bartowski
Q5_K_S4.95 GiB5,313,013,9525.248openbmb
Q5_04.95 GiB5,313,013,9525.248openbmb
Q5_K_S4.95 GiB5,313,014,4005.248bartowski
Q5_K5.07 GiB5,442,668,7365.376openbmb
Q5_K_M5.07 GiB5,442,668,7365.376openbmb
Q5_K_M5.07 GiB5,442,668,9285.376lmstudio-community
Q5_K_M5.07 GiB5,442,669,1845.376bartowski
Q5_15.36 GiB5,754,817,6645.684openbmb
Q5_K_L5.38 GiB5,778,154,3685.707bartowski
Q6_K5.82 GiB6,251,846,8486.175openbmb
Q6_K5.82 GiB6,251,847,0406.175lmstudio-community
Q6_K5.82 GiB6,251,847,2966.175bartowski

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.24 GiB. The real file is 4.36 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does MiniCPM-V-2_6 need?
Q4_K_M is exactly 4,681,089,344 bytes (4.36 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of MiniCPM-V-2_6 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.