Zaynoid · text

Qwen3.5-9B-MedHELM-Reasoning-v1

Zaynoid/Qwen3.5-9B-MedHELM-Reasoning-v1

Qwen3.5-9B-MedHELM-Reasoning-v1 at Q4_K_M is exactly 5,780,091,136 bytes (5.38 GiB / 5.78 GB) — an effective 5.028 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
9.2B
Architecture
qwen35
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q2_K3.65 GiB3,914,969,3443.405mradermacher
I1-Q2_K3.65 GiB3,914,969,6323.405mradermacher
Q3_K_S4.06 GiB4,364,022,0163.796mradermacher
I1-Q3_K_S4.06 GiB4,364,022,3043.796mradermacher
I1-IQ3_S4.17 GiB4,475,990,5603.893mradermacher
I1-IQ3_M4.21 GiB4,522,783,2643.934mradermacher
Q3_K_M4.41 GiB4,737,609,9844.121mradermacher
I1-Q3_K_M4.41 GiB4,737,610,2724.121mradermacher
Q3_K_L4.70 GiB5,048,512,7684.391mradermacher
I1-Q3_K_L4.70 GiB5,048,513,0564.391mradermacher
I1-IQ4_XS4.96 GiB5,326,418,4644.633mradermacher
IQ4_XS4.99 GiB5,357,875,4564.660mradermacher
I1-Q4_05.09 GiB5,462,864,4164.752mradermacher
Q4_K_S5.11 GiB5,488,554,2404.774mradermacher
I1-Q4_K_S5.11 GiB5,488,554,5284.774mradermacher
I1-IQ4_NL5.17 GiB5,555,663,3924.832mradermacher
Q4_K_M5.38 GiB5,780,091,1365.028mradermacher
I1-Q4_K_M5.38 GiB5,780,091,4245.028mradermacher
I1-Q4_15.55 GiB5,961,462,3045.186mradermacher
Q5_K_S6.03 GiB6,472,642,8165.630mradermacher
I1-Q5_K_S6.03 GiB6,472,643,1045.630mradermacher
Q5_K_M6.19 GiB6,642,544,8965.778mradermacher
I1-Q5_K_M6.19 GiB6,642,545,1845.778mradermacher
Q6_K7.04 GiB7,558,902,0166.575mradermacher
I1-Q6_K7.04 GiB7,558,902,3046.575mradermacher
Q8_09.11 GiB9,786,061,0568.512mradermacher
F1617.14 GiB18,407,321,85616.011mradermacher

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 4.82 GiB. The real file is 5.38 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does Qwen3.5-9B-MedHELM-Reasoning-v1 need?
Q4_K_M is exactly 5,780,091,136 bytes (5.38 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of Qwen3.5-9B-MedHELM-Reasoning-v1 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.