Qwen3-Coder-Next-Opus-4.6-Reasoning-Distilled Q3_K
This file is exactly 38,278,053,760 bytes — 35.65 GiB / 38.28 GB.
From the file· 1 file(s), summed
Get it
35.65 GiB · 1 file
llama.cpp
llama-cli -hf Momix-44/Huihui-Qwen3-Coder-Next-Opus-4.6-Reasoning-Distilled-abliterated:Q3_K
Downloads and runs in one step, resolving the quantization by name.
Hugging Face CLI
hf download Momix-44/Huihui-Qwen3-Coder-Next-Opus-4.6-Reasoning-Distilled-abliterated Huihui-Qwen3-Coder-Next-Opus-4.6-Reasoning-Distilled-abliterated-Q3_K.gguf
Direct download
Straight from the Hugging Face CDN — we host nothing and earn nothing from this. Verify what you received against the exact byte count above; a size mismatch is the usual cause of a file that will not load.
Size
35.65 GiB
38.28 GB
Effective bpw
—
from real bytes ÷ params
Tensors
—
Header
—
GGUF metadata
Fit by accelerator
| Accelerator | Memory○ | Total @4K◐ | Total @32K◐ | Fits◐ | tok/s @4K◐ |
|---|---|---|---|---|---|
| Arc A310 4GB | 4 GB | 36.53 GiB | 37.19 GiB | no | — |
| Arc A350M 4GB | 4 GB | 36.53 GiB | 37.19 GiB | no | — |
| Arc A370M 4GB | 4 GB | 36.53 GiB | 37.19 GiB | no | — |
| Arc A530M 4GB | 4 GB | 36.53 GiB | 37.19 GiB | no | — |
| Arc Pro A30M 4GB | 4 GB | 36.53 GiB | 37.19 GiB | no | — |
| Radeon Pro W6400 | 4 GB | 36.63 GiB | 37.29 GiB | no | — |
| Radeon RX 6400 | 4 GB | 36.63 GiB | 37.29 GiB | no | — |
| Radeon RX 6500 XT | 4 GB | 36.63 GiB | 37.29 GiB | no | — |
| RTX A400 | 4 GB | 36.73 GiB | 37.39 GiB | no | — |
| Arc A380 6GB | 6 GB | 36.53 GiB | 37.19 GiB | no | — |
| Arc Pro A40 6GB | 6 GB | 36.53 GiB | 37.19 GiB | no | — |
| Arc Pro A50 6GB | 6 GB | 36.53 GiB | 37.19 GiB | no | — |
| GeForce RTX 2060 | 6 GB | 36.53 GiB | 37.19 GiB | no | — |
| GeForce RTX 3050 | 6 GB | 36.53 GiB | 37.19 GiB | no | — |
| GeForce RTX 3060 OEM | 6 GB | 36.53 GiB | 37.19 GiB | no | — |
| RTX A2000 | 6 GB | 36.73 GiB | 37.39 GiB | no | — |
| Apple M1 | 8 GB | 36.28 GiB | 36.94 GiB | no | — |
| Apple M2 | 8 GB | 36.28 GiB | 36.94 GiB | no | — |
| Apple M3 | 8 GB | 36.28 GiB | 36.94 GiB | no | — |
| Apple M4 | 8 GB | 36.28 GiB | 36.94 GiB | no | — |
| Arc A530M 8GB | 8 GB | 36.53 GiB | 37.19 GiB | no | — |
| Arc A550M 8GB | 8 GB | 36.53 GiB | 37.19 GiB | no | — |
| Arc A570M 8GB | 8 GB | 36.53 GiB | 37.19 GiB | no | — |
| Arc A580 8GB | 8 GB | 36.53 GiB | 37.19 GiB | no | — |
Memory at context
| Context | Weights● | KV (f16)● | KV (q8_0)● | Working buffer◐ | Total (f16)◐ |
|---|---|---|---|---|---|
| 4,096 | 35.65 GiB | 0.09 GiB | 0.05 GiB | 0.29 GiB | 36.03 GiB |
| 8,192 | 35.65 GiB | 0.19 GiB | 0.10 GiB | 0.29 GiB | 36.13 GiB |
| 16,384 | 35.65 GiB | 0.38 GiB | 0.20 GiB | 0.29 GiB | 36.31 GiB |
| 32,768 | 35.65 GiB | 0.75 GiB | 0.40 GiB | 0.29 GiB | 36.69 GiB |
| 65,536 | 35.65 GiB | 1.50 GiB | 0.80 GiB | 0.29 GiB | 37.44 GiB |
| 131,072 | 35.65 GiB | 3.00 GiB | 1.59 GiB | 0.29 GiB | 38.94 GiB |
Quantizing the KV cache to q8_0 is a roughly 2× lever on the dominant term at long context, and it is the single most useful setting most local users never touch. Totals here exclude the allocator reserve your driver takes, which is hardware-specific — the per-accelerator table above includes it.