meta-llama · text

Meta-Llama-3-70B-Instruct

meta-llama/Meta-Llama-3-70B-Instruct

Meta-Llama-3-70B-Instruct at Q4_K_M is exactly 42,520,393,120 bytes (39.60 GiB / 42.52 GB) — an effective 4.821 bits per weight, not the nominal 4. Its KV cache at 32K is 10.00 GiB.

From the file· summed from 1 file(s)From the file· KV from mirror (mirror:NousResearch/Meta-Llama-3-70B-Instruct)
Parameters
70.6B
Architecture
llama
80 layers
Context
8,192
native (config.json)
License
other

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ1_S14.29 GiB15,343,482,2721.740qwp4w3hyb
IQ1_M15.60 GiB16,751,195,5521.899lmstudio-community
IQ2_XXS17.79 GiB19,097,384,3522.165qwp4w3hyb
IQ2_XS19.69 GiB21,142,107,5522.397lmstudio-community
IQ2_XS19.69 GiB21,142,107,5522.397qwp4w3hyb
IQ2_S20.71 GiB22,242,342,3042.522qwp4w3hyb
IQ2_M22.46 GiB24,119,293,3442.735qwp4w3hyb
Q2_K24.56 GiB26,375,108,2242.991LiteLLMs
IQ3_XXS25.58 GiB27,469,493,6643.115qwp4w3hyb
IQ3_XS27.29 GiB29,307,729,3123.323qwp4w3hyb
IQ3_S28.79 GiB30,912,050,5923.505qwp4w3hyb
Q3_K_S28.79 GiB30,912,050,8163.505LiteLLMs
IQ3_M29.74 GiB31,937,033,6323.621qwp4w3hyb
Q3_K_M3 shards31.92 GiB34,275,312,9283.886LiteLLMs
Q3_K_L3 shards34.60 GiB37,148,411,1684.212LiteLLMs
IQ4_XS35.30 GiB37,902,661,0244.298qwp4w3hyb
Q4_03 shards37.23 GiB39,977,551,1364.533LiteLLMs
IQ4_NL37.30 GiB40,053,618,0804.542qwp4w3hyb
Q4_K_S37.58 GiB40,347,219,3604.575qwp4w3hyb
Q4_K_S3 shards37.58 GiB40,355,038,4964.576LiteLLMs
Q4_K_M39.60 GiB42,520,393,1204.821lmstudio-community
Q4_K_M39.60 GiB42,520,393,1204.821qwp4w3hyb
Q4_K_M3 shards39.61 GiB42,528,212,2564.822LiteLLMs
Q4_13 shards41.28 GiB44,321,408,2885.026LiteLLMs
Q5_K_S45.32 GiB48,657,446,3045.517qwp4w3hyb
Q5_03 shards45.32 GiB48,665,265,4085.518LiteLLMs
Q5_K_S3 shards45.32 GiB48,665,265,4085.518LiteLLMs
Q5_K_M46.52 GiB49,949,816,2245.664qwp4w3hyb
Q5_K_M3 shards46.53 GiB49,957,635,3605.665LiteLLMs
Q5_13 shards49.37 GiB53,009,122,5926.011LiteLLMs
Q6_K3 shards53.92 GiB57,895,961,8886.565LiteLLMs
Q8_04 shards69.83 GiB74,982,868,3528.502LiteLLMs

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0961.25 GiB1.25 GiB80 / 0 / 0
8,1922.50 GiB2.50 GiB80 / 0 / 0
16,3845.00 GiB5.00 GiB80 / 0 / 0
32,76810.00 GiB10.00 GiB80 / 0 / 0
65,53620.00 GiB20.00 GiB80 / 0 / 0
131,07240.00 GiB40.00 GiB80 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 36.96 GiB. The real file is 39.60 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from mirror:NousResearch/Meta-Llama-3-70B-Instruct
Layers
80
Attention heads
64
KV heads
8
Head dim
128
Hidden size
8192
Vocab
128,256
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Meta-Llama-3-70B-Instruct need?
Q4_K_M is exactly 42,520,393,120 bytes (39.60 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Meta-Llama-3-70B-Instruct's KV cache?
10.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Meta-Llama-3-70B-Instruct should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.