ZERO-POINT-AI · text

MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_base

ZERO-POINT-AI/MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_base

MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_base at Q4_K_M is exactly 5,106,435,872 bytes (4.76 GiB / 5.11 GB) — an effective 5.050 bits per weight, not the nominal 4. Its KV cache at 32K is 0.88 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
8.1B
Architecture
qwen35
28 layers
Context
262,144
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S2.35 GiB2,526,965,7922.499mradermacher
I1-IQ1_M2.46 GiB2,645,354,5282.616mradermacher
I1-IQ2_XXS2.65 GiB2,842,669,0882.811mradermacher
I1-IQ2_XS2.80 GiB3,004,190,7522.971mradermacher
I1-IQ2_S2.92 GiB3,139,342,3683.105mradermacher
I1-IQ2_M3.07 GiB3,297,194,0163.261mradermacher
I1-Q2_K_S3.14 GiB3,376,568,3523.340mradermacher
Q2_K3.26 GiB3,496,236,8323.458mradermacher
I1-Q2_K3.26 GiB3,496,237,0883.458mradermacher
I1-IQ3_XXS3.34 GiB3,589,304,3523.550mradermacher
I1-IQ3_XS3.61 GiB3,873,284,1283.831mradermacher
Q3_K_S3.62 GiB3,887,275,8083.845mradermacher
I1-Q3_K_S3.62 GiB3,887,276,0643.845mradermacher
I1-IQ3_S3.71 GiB3,984,760,8643.941mradermacher
I1-IQ3_M3.74 GiB4,020,412,4483.976mradermacher
Q3_K_M3.91 GiB4,202,209,0564.156mradermacher
I1-Q3_K_M3.91 GiB4,202,209,3124.156mradermacher
Q3_K_L4.16 GiB4,470,120,2244.421mradermacher
I1-Q3_K_L4.16 GiB4,470,120,4804.421mradermacher
I1-IQ4_XS4.40 GiB4,720,093,2164.668mradermacher
IQ4_XS4.42 GiB4,743,685,9204.692mradermacher
I1-Q4_04.50 GiB4,835,805,2164.783mradermacher
Q4_K_S4.52 GiB4,858,349,3444.805mradermacher
I1-Q4_K_S4.52 GiB4,858,349,6004.805mradermacher
I1-IQ4_NL4.58 GiB4,918,118,4324.864mradermacher
Q4_K_M4.76 GiB5,106,435,8725.050mradermacher
I1-Q4_K_M4.76 GiB5,106,436,1285.050mradermacher
I1-Q4_14.91 GiB5,268,293,6645.210mradermacher
Q5_K_S5.32 GiB5,710,219,0405.647mradermacher
I1-Q5_K_S5.32 GiB5,710,219,2965.647mradermacher
Q5_K_M5.45 GiB5,854,496,5445.790mradermacher
I1-Q5_K_M5.45 GiB5,854,496,8005.790mradermacher
Q6_K6.19 GiB6,649,311,0086.576mradermacher
I1-Q6_K6.19 GiB6,649,311,2646.576mradermacher
Q8_08.02 GiB8,608,106,2728.514mradermacher
F1615.08 GiB16,190,539,55216.013mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.11 GiB0.44 GiB4.00×7 / 0 / 21
8,1920.22 GiB0.88 GiB4.00×7 / 0 / 21
16,3840.44 GiB1.75 GiB4.00×7 / 0 / 21
32,7680.88 GiB3.50 GiB4.00×7 / 0 / 21
65,5361.75 GiB7.00 GiB4.00×7 / 0 / 21
131,0723.50 GiB14.00 GiB4.00×7 / 0 / 21

21 of 28 layers use linear attention, which keeps a fixed-size recurrent state instead of a per-token cache. Those layers do not grow with context at all — treating them as ordinary attention, as a flat formula does, overstates this model's cache by roughly 4.0× at long context.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.24 GiB. The real file is 4.76 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
28
Attention heads
16
KV heads
4
Head dim
256
Hidden size
4096
Vocab
248,320
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_base need?
Q4_K_M is exactly 5,106,435,872 bytes (4.76 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_base's KV cache?
0.88 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of MARTHA-LXVII.8B_QWEN-3.5-3.6_prune_9b-3.6_base should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.