xtuner · vision language

llava-llama-3-8b-v1_1-transformers

xtuner/llava-llama-3-8b-v1_1-transformers

llava-llama-3-8b-v1_1-transformers at Q4_K_M is exactly 4,921,098,752 bytes (4.58 GiB / 4.92 GB) — an effective 4.712 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
8.4B
Architecture
llama
null layers
Context
8,192
native (config.json)
License

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q3_K_S3.41 GiB3,664,828,9283.509city96
Q3_K_M3.74 GiB4,019,247,6163.848city96
IQ4_XS4.14 GiB4,448,018,9444.259city96
Q4_04.36 GiB4,676,256,2564.477city96
IQ4_NL4.36 GiB4,678,353,4084.479city96
Q4_K_S4.37 GiB4,693,033,4724.494city96
Q4_K_M4.58 GiB4,921,098,7524.712city96
Q4_14.78 GiB5,130,633,7284.912city96
Q5_K_S5.22 GiB5,599,691,2645.362city96
Q5_05.23 GiB5,614,371,3285.376city96
Q5_K_M5.34 GiB5,733,384,7045.490city96
Q5_15.65 GiB6,068,748,8005.811city96
Q6_K6.14 GiB6,596,438,3366.316city96
Q8_07.95 GiB8,541,329,7288.178city96
F1614.97 GiB16,069,941,56815.387city96

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.38 GiB. The real file is 4.58 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
Attention heads
KV heads
8
Head dim
Hidden size
Vocab
128,320
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does llava-llama-3-8b-v1_1-transformers need?
Q4_K_M is exactly 4,921,098,752 bytes (4.58 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of llava-llama-3-8b-v1_1-transformers should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.