Submit a model
Drop a HuggingFace reference and we will benchmark it on our own hardware β quality on the private set, plus a full serving perf sweep. No account needed.
what we require
- safetensors weights. We do not load `.bin` checkpoints β they are pickles, and loading one is remote code execution on our hardware.
- Ungated repo. We cannot download what we do not have access to.
- Servable size. It has to fit the box, at some quantization we can actually run.
what you get back
- A public score card: aggregate quality plus the full concurrency sweep, on named hardware with the serving flags recorded.
- Per-sample breakdowns stay with signed-in accounts. Detailed per-sample feedback is how a private eval set gets reverse-engineered, so anonymous results are aggregate only.
- One machine, one queue. Big models wait.