nvidia · audio asr

canary-1b-v2

nvidia/canary-1b-v2

canary-1b-v2 at Q4_K_M is exactly 735,476,448 bytes (0.68 GiB / 0.74 GB) — an effective 6.004 bits per weight, not the nominal 4.

From the file· summed from 1 file(s)
Parameters
980M
Architecture
canary
Context
native (config.json)
License
cc-by-4.0

Shipped quantizations

exact bytes, summed from published files
QuantSizeExact bytesEffective bpwTensorsPublisher
Q4_K0.37 GiB392,167,0403.201cstr
Q5_00.43 GiB460,848,8963.762cstr
Q4_K0.57 GiB610,746,0164.985cstr
Q8_00.62 GiB666,894,4645.444712cstr
Q5_00.67 GiB720,322,2085.880cstr
Q4_K_M0.68 GiB735,476,4486.0041476handy-computer
Q5_K_M0.78 GiB836,664,0326.8301476handy-computer
Q6_K0.87 GiB931,986,1447.6081476handy-computer
Q8_00.98 GiB1,049,050,7848.563cstr
Q8_01.07 GiB1,144,290,0169.3411476handy-computer
F161.83 GiB1,966,111,45616.0491476handy-computer
F323.65 GiB3,920,657,12032.004handy-computer

Measured

published by a third party, attributed below
MetricValueWhat it means
rtf1821.4
RTFx1821.4Higher is better — audio seconds processed per second of compute.
Word error rate14.71%Lower is better — the share of words transcribed incorrectly.
Word error rate5.93%Lower is better — the share of words transcribed incorrectly.
Word error rate9.20%Lower is better — the share of words transcribed incorrectly.
Word error rate4.39%Lower is better — the share of words transcribed incorrectly.
Word error rate2.03%Lower is better — the share of words transcribed incorrectly.
Word error rate3.07%Lower is better — the share of words transcribed incorrectly.
Word error rate1.76%Lower is better — the share of words transcribed incorrectly.
Word error rate9.12%Lower is better — the share of words transcribed incorrectly.
Word error rate13.02%Lower is better — the share of words transcribed incorrectly.
Word error rate11.35%Lower is better — the share of words transcribed incorrectly.
Word error rate6.39%Lower is better — the share of words transcribed incorrectly.
Benchmarked· by open-asr-leaderboard-english-short-latest

RTFx measured by the Open ASR Leaderboard on a single datacenter GPU at a large batch size. It ranks models against each other; it says nothing about throughput on consumer hardware. We reproduce these figures with attribution; they are not ours and we have not verified the runs. Source: open-asr-leaderboard-english-short-latest.

Compare with

same modality, comparable size

Will it run on your card?

full quant x context sweep

Why other calculators give a different number

A parameters × bits ÷ 8 estimate puts Q4_K_M at roughly 0.51 GiB. The real file is 0.68 GiB, because a quantization is a mixture and some tensors are always kept at higher precision.

Architecture

Architecture unavailable — this repository is gated and no ungated mirror was found. Exact file sizes above are still authoritative; only the KV math needs the config.

Questions people ask

How much VRAM does canary-1b-v2 need?
Q4_K_M is exactly 735,476,448 bytes (0.68 GiB) in weights. Add the KV cache, which depends on your context length, plus roughly half a gigabyte of runtime overhead.
Which quantization of canary-1b-v2 should I use?
Q4_K_M is the usual default. Pick the largest quantization that fits your card at the context you actually need — the table above gives exact sizes for every one published.