Apertus-8B-Instruct-2509 Q2_K
This file is exactly 3,287,889,152 bytes — 3.06 GiB / 3.29 GB at an effective 3.266 bits per weight. The nominal rate for Q2_K is lower; mixed-precision tensors make the real figure higher, always.
Get it
llama-cli -hf unsloth/Apertus-8B-Instruct-2509-GGUF:Q2_K
Downloads and runs in one step, resolving the quantization by name.
hf download unsloth/Apertus-8B-Instruct-2509-GGUF Apertus-8B-Instruct-2509-Q2_K.gguf
Straight from the Hugging Face CDN — we host nothing and earn nothing from this. Verify what you received against the exact byte count above; a size mismatch is the usual cause of a file that will not load.
Fit by accelerator
| Accelerator | Memory○ | Total @4K◐ | Total @32K◐ | Fits◐ | tok/s @4K◐ |
|---|---|---|---|---|---|
| Arc A310 4GB | 4 GB | 4.43 GiB | 7.93 GiB | no | — |
| Arc A350M 4GB | 4 GB | 4.43 GiB | 7.93 GiB | no | — |
| Arc A370M 4GB | 4 GB | 4.43 GiB | 7.93 GiB | no | — |
| Arc A530M 4GB | 4 GB | 4.43 GiB | 7.93 GiB | no | — |
| Arc Pro A30M 4GB | 4 GB | 4.43 GiB | 7.93 GiB | no | — |
| Radeon Pro W6400 | 4 GB | 4.53 GiB | 8.03 GiB | no | — |
| Radeon RX 6400 | 4 GB | 4.53 GiB | 8.03 GiB | no | — |
| Radeon RX 6500 XT | 4 GB | 4.53 GiB | 8.03 GiB | no | — |
| RTX A400 | 4 GB | 4.63 GiB | 8.13 GiB | no | — |
| Arc A380 6GB | 6 GB | 4.43 GiB | 7.93 GiB | yes4K only | 28±30% |
| Arc Pro A40 6GB | 6 GB | 4.43 GiB | 7.93 GiB | yes4K only | 28±30% |
| Arc Pro A50 6GB | 6 GB | 4.43 GiB | 7.93 GiB | yes4K only | 28±30% |
| GeForce RTX 2060 | 6 GB | 4.43 GiB | 7.93 GiB | yes4K only | 66±12.9% |
| GeForce RTX 3050 | 6 GB | 4.43 GiB | 7.93 GiB | yes4K only | 34±12.9% |
| GeForce RTX 3060 OEM | 6 GB | 4.43 GiB | 7.93 GiB | yes4K only | 66±12.9% |
| RTX A2000 | 6 GB | 4.63 GiB | 8.13 GiB | yes4K only | 45±22% |
| Apple M1 | 8 GB | 4.18 GiB | 7.68 GiB | yes4K only | 15±8.3% |
| Apple M2 | 8 GB | 4.18 GiB | 7.68 GiB | yes4K only | 22±8.3% |
| Apple M3 | 8 GB | 4.18 GiB | 7.68 GiB | yes4K only | 22±8.3% |
| Apple M4 | 8 GB | 4.18 GiB | 7.68 GiB | yes4K only | 25±8.3% |
| Arc A530M 8GB | 8 GB | 4.43 GiB | 7.93 GiB | yes4K only | 33±30% |
| Arc A550M 8GB | 8 GB | 4.43 GiB | 7.93 GiB | yes4K only | 33±30% |
| Arc A570M 8GB | 8 GB | 4.43 GiB | 7.93 GiB | yes4K only | 33±30% |
| Arc A580 8GB | 8 GB | 4.43 GiB | 7.93 GiB | yes4K only | 69±30% |
Memory at context
| Context | Weights● | KV (f16)● | KV (q8_0)● | Working buffer◐ | Total (f16)◐ |
|---|---|---|---|---|---|
| 4,096 | 3.06 GiB | 0.50 GiB | 0.27 GiB | 0.37 GiB | 3.93 GiB |
| 8,192 | 3.06 GiB | 1.00 GiB | 0.53 GiB | 0.37 GiB | 4.43 GiB |
| 16,384 | 3.06 GiB | 2.00 GiB | 1.06 GiB | 0.37 GiB | 5.43 GiB |
| 32,768 | 3.06 GiB | 4.00 GiB | 2.13 GiB | 0.37 GiB | 7.43 GiB |
| 65,536 | 3.06 GiB | 8.00 GiB | 4.25 GiB | 0.37 GiB | 11.43 GiB |
| 131,072 | 3.06 GiB | 16.00 GiB | 8.50 GiB | 0.37 GiB | 19.43 GiB |
Quantizing the KV cache to q8_0 is a roughly 2× lever on the dominant term at long context, and it is the single most useful setting most local users never touch. Totals here exclude the allocator reserve your driver takes, which is hardware-specific — the per-accelerator table above includes it.