p-e-w · text

Llama-3.1-8B-Instruct-heretic

p-e-w/Llama-3.1-8B-Instruct-heretic

Llama-3.1-8B-Instruct-heretic at Q4_K_M is exactly 4,920,739,808 bytes (4.58 GiB / 4.92 GB) — an effective 4.902 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
8.0B
Architecture
llama
Context
native (config.json)
License
llama3.1

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
IQ2_M2.75 GiB2,948,286,4322.937bartowski
Q2_K2.96 GiB3,179,136,9923.167bartowski
IQ3_XXS3.05 GiB3,274,917,8563.263bartowski
IQ3_XS3.28 GiB3,518,752,7363.506bartowski
Q3_K_S3.41 GiB3,664,504,8003.651bartowski
Q2_K_L3.44 GiB3,692,160,9923.678bartowski
IQ3_M3.52 GiB3,784,828,8963.771bartowski
Q3_K_M3.74 GiB4,018,923,4884.004bartowski
I1-IQ1_S2 shards3.76 GiB4,039,267,1364.024mradermacher
Q3_K_L4.03 GiB4,321,961,9524.306bartowski
I1-IQ1_M2 shards4.03 GiB4,323,955,5204.308mradermacher
IQ4_XS4.14 GiB4,447,668,1924.431bartowski
Q4_04.35 GiB4,675,897,3124.658bartowski
IQ4_NL4.36 GiB4,677,994,4644.660bartowski
Q4_K_S4.37 GiB4,692,674,5284.675bartowski
I1-IQ2_XXS2 shards4.47 GiB4,798,436,1604.780mradermacher
Q4_K_M4.58 GiB4,920,739,8084.902bartowski
Q4_14.78 GiB5,130,258,4005.111bartowski
I1-IQ2_XS2 shards4.85 GiB5,211,575,1045.192mradermacher
Q4_K_L4.95 GiB5,310,638,0485.291bartowski
I1-IQ2_S2 shards5.14 GiB5,516,989,2485.496mradermacher
Q5_K_S5.21 GiB5,599,299,5525.578bartowski
Q5_K_M5.34 GiB5,732,992,9925.711bartowski
I1-IQ2_M2 shards5.49 GiB5,896,573,7605.874mradermacher
I1-Q2_K_S2 shards5.57 GiB5,977,641,7925.955mradermacher
Q5_K_L5.64 GiB6,057,224,1606.034bartowski
Q2_K2 shards5.92 GiB6,358,274,3366.334mradermacher
I1-Q2_K2 shards5.92 GiB6,358,274,8806.334mradermacher
I1-IQ3_XXS2 shards6.10 GiB6,549,836,6086.525mradermacher
Q6_K6.14 GiB6,596,012,0006.571bartowski
Q6_K_L6.38 GiB6,850,471,9046.825bartowski
I1-IQ3_XS2 shards6.55 GiB7,037,506,3687.011mradermacher
Q3_K_S2 shards6.83 GiB7,329,009,9527.301mradermacher
I1-Q3_K_S2 shards6.83 GiB7,329,010,4967.301mradermacher
I1-IQ3_S2 shards6.86 GiB7,364,662,0807.337mradermacher
I1-IQ3_M2 shards7.05 GiB7,569,658,6887.541mradermacher
Q3_K_M2 shards7.49 GiB8,037,847,3288.008mradermacher
I1-Q3_K_M2 shards7.49 GiB8,037,847,8728.008mradermacher
Q8_07.95 GiB8,540,776,4168.509bartowski
Q3_K_L2 shards8.05 GiB8,643,924,2568.611mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.21 GiB. The real file is 4.58 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Llama-3.1-8B-Instruct-heretic need?
Q4_K_M is exactly 4,920,739,808 bytes (4.58 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Llama-3.1-8B-Instruct-heretic should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.