ServiceNow-AI · vision language

Apriel-1.6-15b-Thinker

ServiceNow-AI/Apriel-1.6-15b-Thinker

Apriel-1.6-15b-Thinker at Q4_K_M is exactly 8,785,478,144 bytes (8.18 GiB / 8.79 GB) — an effective 4.729 bits per weight, not the nominal 4. Its KV cache at 32K is 6.00 GiB.

From the file· summed from 1 file(s)From the file· KV per layer
Parameters
14.9B
Architecture
llama
48 layers
Context
262,400
native (config.json)
License
mit

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_S3.22 GiB3,461,170,0801.863mradermacher
I1-IQ1_M3.47 GiB3,728,065,4402.006mradermacher
I1-IQ2_XXS3.89 GiB4,172,891,0402.246mradermacher
IQ2_XS4.25 GiB4,560,208,3842.454bartowski
I1-IQ2_XS4.25 GiB4,560,208,8002.454mradermacher
IQ2_S4.48 GiB4,814,651,9042.591bartowski
I1-IQ2_S4.48 GiB4,814,652,3202.591mradermacher
IQ2_M4.82 GiB5,170,512,3842.783bartowski
I1-IQ2_M4.82 GiB5,170,512,8002.783mradermacher
I1-Q2_K_S4.88 GiB5,236,704,1602.818mradermacher
Q2_K5.21 GiB5,593,547,2643.010bartowski
Q2_K5.21 GiB5,593,547,4243.010mradermacher
I1-Q2_K5.21 GiB5,593,547,6803.010mradermacher
IQ3_XXS5.39 GiB5,782,946,3043.112bartowski
I1-IQ3_XXS5.39 GiB5,782,946,7203.112mradermacher
IQ3_XS5.77 GiB6,198,444,5443.336bartowski
I1-IQ3_XS5.77 GiB6,198,444,9603.336mradermacher
Q2_K_L5.82 GiB6,248,907,2643.363bartowski
Q3_K_S6.03 GiB6,471,729,6643.483bartowski
Q3_K_S6.03 GiB6,471,729,8243.483mradermacher
I1-Q3_K_S6.03 GiB6,471,730,0803.483mradermacher
I1-IQ3_S6.06 GiB6,505,153,4403.501mradermacher
IQ3_M6.24 GiB6,697,337,3443.605bartowski
I1-IQ3_M6.24 GiB6,697,337,7603.605mradermacher
Q3_K_M6.65 GiB7,135,609,3443.841bartowski
Q3_K_M6.65 GiB7,135,609,5043.841mradermacher
I1-Q3_K_M6.65 GiB7,135,609,7603.841mradermacher
Q3_K_L7.18 GiB7,704,461,8244.147bartowski
Q3_K_L7.18 GiB7,704,461,9844.147mradermacher
I1-Q3_K_L7.18 GiB7,704,462,2404.147mradermacher
IQ4_XS7.37 GiB7,908,278,7844.256bartowski
I1-IQ4_XS7.37 GiB7,908,279,2004.256mradermacher
IQ4_XS7.43 GiB7,977,091,7444.293mradermacher
Q4_07.75 GiB8,326,398,4644.481bartowski
I1-Q4_07.75 GiB8,326,398,8804.481mradermacher
IQ4_NL7.76 GiB8,330,330,6244.484bartowski
I1-IQ4_NL7.76 GiB8,330,331,0404.484mradermacher
Q4_K_S7.78 GiB8,356,545,0244.498bartowski
Q4_K_S7.78 GiB8,356,545,1844.498mradermacher
I1-Q4_K_S7.78 GiB8,356,545,4404.498mradermacher

KV cache by context

computed per layer
ContextKV cache (f16)Flat formulaOverstated byFull / windowed / recurrent
4,0960.75 GiB0.75 GiB48 / 0 / 0
8,1921.50 GiB1.50 GiB48 / 0 / 0
16,3843.00 GiB3.00 GiB48 / 0 / 0
32,7686.00 GiB6.00 GiB48 / 0 / 0
65,53612.00 GiB12.00 GiB48 / 0 / 0
131,07224.00 GiB24.00 GiB48 / 0 / 0

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 7.79 GiB. The real file is 8.18 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

from config.json
Layers
48
Attention heads
32
KV heads
8
Head dim
128
Hidden size
5120
Vocab
131,072
Sliding window
none
SWA period
MLA
no
Experts
Experts per token
use_sliding_window

Questions people ask

How much VRAM does Apriel-1.6-15b-Thinker need?
Q4_K_M is exactly 8,785,478,144 bytes (8.18 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
How large is Apriel-1.6-15b-Thinker's KV cache?
6.00 GiB at 32K context with an f16 cache, computed per layer. Quantizing the cache to q8_0 roughly halves it, which is often the difference between a context length fitting and not.
Which quantization of Apriel-1.6-15b-Thinker should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.