Status

What is in the dataset right now, and what it cannot yet tell you.

Base models
2,493
Quantizations
67,146
Accelerators
239
Fit counts
478
Measured rows
57,390

Coverage

last ingest 2026-07-28 15:16
FieldCountCoverage
architecture resolved2,16787%
from an ungated mirror502%
effective bpw computed66,65599%
tensor histogram parsed25610%

Measured data

117 accelerators covered

Reproduced from third-party benchmark publishers with attribution, never relicensed. None of it was measured by us, and we say so wherever it appears. It exists so that our own modeled figures can eventually be scored against something real rather than merely banded.

By modality

audio asr 52image 47video 40embedding 27audio tts 25vision language 216text 2086

Known limitations

No first-party measured performance
We now reproduce third-party measured benchmarks with attribution — image generation throughput per GPU, and speech recognition accuracy and speed. But every figure WE compute is still modeled from memory bandwidth with a published error band. Nothing on this site has been measured by us.
Measured data is not redistributable
The sources we reproduce publish no licence of their own. We display and attribute them, and exclude them from our CC-BY bulk dump, which therefore contains only data we derived ourselves.
Image throughput is community-submitted, not controlled
The image-generation figures per GPU aggregate thousands of community runs across different models, resolutions and step counts. Read the spread, not the median alone — it is not one controlled configuration.
Video has no throughput data at all
Component sizes are exact. Seconds per clip on consumer hardware is unpublished anywhere credible, so we show nothing rather than a guess.
The catalog is a slice, not the whole Hub
We index the most-downloaded base models. The full deduplicated set is roughly 1,500–3,000.
Gated models may lack architecture
Exact file sizes are always available, even for gated repositories. Architecture sometimes has to come from an ungated mirror, which we flag with reduced confidence, or is unavailable entirely.
Compute buffer is modeled with constants we chose
There is no closed form for the runtime's working buffers. Ours scales with the terms that should drive it — feed-forward width, micro-batch, and context without flash attention — but its constants are not fitted against measured allocations. It is the widest error band on the site and the reason a total carries one when its parts do not.