piotreknow02 · text

GPT-OSS-Cybersecurity-20B-Merged-heretic

piotreknow02/GPT-OSS-Cybersecurity-20B-Merged-heretic

GPT-OSS-Cybersecurity-20B-Merged-heretic at Q4_K_M is exactly 15,805,138,496 bytes (14.72 GiB / 15.81 GB) — an effective 6.045 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
20.9B
Architecture
gpt-oss
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
I1-IQ1_M11.19 GiB12,016,855,9044.596mradermacher
I1-IQ2_XXS11.19 GiB12,016,855,9044.596mradermacher
I1-IQ1_S11.19 GiB12,016,855,9044.596mradermacher
I1-IQ2_XS11.20 GiB12,025,703,2644.600mradermacher
Q3_K_S11.23 GiB12,061,092,4164.613mradermacher
I1-Q3_K_S11.23 GiB12,061,092,7044.613mradermacher
Q2_K11.24 GiB12,065,516,0964.615mradermacher
I1-IQ3_XXS11.24 GiB12,065,516,3844.615mradermacher
I1-IQ3_XS11.24 GiB12,065,516,3844.615mradermacher
I1-Q2_K11.24 GiB12,065,516,3844.615mradermacher
I1-IQ2_M11.24 GiB12,065,516,3844.615mradermacher
I1-IQ3_S11.24 GiB12,065,516,3844.615mradermacher
I1-IQ2_S11.24 GiB12,065,516,3844.615mradermacher
I1-IQ4_XS11.27 GiB12,096,482,1444.627mradermacher
I1-Q2_K_S11.30 GiB12,136,295,2644.642mradermacher
I1-Q4_011.31 GiB12,148,460,3844.647mradermacher
I1-IQ3_M11.36 GiB12,202,650,4644.668mradermacher
IQ4_XS11.40 GiB12,245,781,0564.684mradermacher
Q3_K_M12.03 GiB12,916,152,8964.941mradermacher
I1-Q3_K_M12.03 GiB12,916,153,1844.941mradermacher
Q3_K_L12.42 GiB13,335,112,2565.101mradermacher
I1-Q3_K_L12.42 GiB13,335,112,5445.101mradermacher
I1-Q4_112.45 GiB13,369,096,5445.114mradermacher
Q4_K_S13.65 GiB14,654,244,4165.605mradermacher
I1-Q4_K_S13.65 GiB14,654,244,7045.605mradermacher
Q4_K_M14.72 GiB15,805,138,4966.045mradermacher
I1-Q4_K_M14.72 GiB15,805,138,7846.045mradermacher
Q5_K_S14.80 GiB15,892,206,6566.079mradermacher
I1-Q5_K_S14.80 GiB15,892,206,9446.079mradermacher
Q5_K_M15.73 GiB16,893,064,2566.462mradermacher
I1-Q5_K_M15.73 GiB16,893,064,5446.462mradermacher
Q6_K20.67 GiB22,193,347,1368.489mradermacher
I1-Q6_K20.67 GiB22,193,347,4248.489mradermacher
Q8_020.73 GiB22,261,914,1768.515mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 10.96 GiB. The real file is 14.72 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does GPT-OSS-Cybersecurity-20B-Merged-heretic need?
Q4_K_M is exactly 15,805,138,496 bytes (14.72 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of GPT-OSS-Cybersecurity-20B-Merged-heretic should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.