CohereLabs · audio asr

cohere-transcribe-03-2026

CohereLabs/cohere-transcribe-03-2026

cohere-transcribe-03-2026 at Q4_K_M is exactly 1,558,162,944 bytes (1.45 GiB / 1.56 GB) — an effective 6.034 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
2.1B
Architecture
cohere_asr
Context
native (config.json)
License
apache-2.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q4_K1.41 GiB1,510,363,0085.849cstr
Q4_K_M1.45 GiB1,558,162,9446.0342103handy-computer
Q5_01.62 GiB1,738,723,2006.733cstr
Q5_K_M1.65 GiB1,770,270,2086.8562103handy-computer
Q5_11.73 GiB1,852,903,2967.176cstr
Q6_K1.84 GiB1,972,524,5447.6392103handy-computer
Q6_K1.85 GiB1,981,355,9047.6732104cstr
Q8_02.25 GiB2,410,655,2329.3352103handy-computer
Q8_02.26 GiB2,423,803,7769.3862104cstr
BF163.82 GiB4,105,263,10415.898handy-computer
F163.82 GiB4,106,644,99215.9032103handy-computer

Measured

published by a third party, attributed below
MetricValueWhat it means
rtf915.6
RTFx915.6Higher is better — audio seconds processed per second of compute.
Word error rate0.96%Lower is better — the share of words transcribed incorrectly.
Word error rate7.88%Lower is better — the share of words transcribed incorrectly.
Word error rate7.80%Lower is better — the share of words transcribed incorrectly.
Word error rate7.93%Lower is better — the share of words transcribed incorrectly.
Word error rate5.58%Lower is better — the share of words transcribed incorrectly.
Word error rate5.38%Lower is better — the share of words transcribed incorrectly.
Word error rate2.74%Lower is better — the share of words transcribed incorrectly.
Word error rate7.01%Lower is better — the share of words transcribed incorrectly.
Word error rate10.38%Lower is better — the share of words transcribed incorrectly.
Word error rate2.04%Lower is better — the share of words transcribed incorrectly.
Word error rate5.20%Lower is better — the share of words transcribed incorrectly.
Benchmarked· by open-asr-leaderboard-english-short-latest

RTFx measured by the Open ASR Leaderboard on a single datacenter GPU at a large batch size. It ranks models against each other; it says nothing about throughput on consumer hardware. We reproduce these figures with attribution; they are not ours and we have not verified the runs. Source: open-asr-leaderboard-english-short-latest.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 1.08 GiB. The real file is 1.45 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does cohere-transcribe-03-2026 need?
Q4_K_M is exactly 1,558,162,944 bytes (1.45 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of cohere-transcribe-03-2026 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.