bench.karti.ai

Does this model earn
its place on our hardware?

Two numbers decide a self-hosting call: whether a model is good enough at the work we actually do, and whether it is fast enough on the box we would run it on. Most leaderboards report only the first. Every card here reports both, measured together, on the same machine.

1 target2 runslast measured 2026-07-28
in production
primary target

Qwen3.6-35B-A3B (NVFP4)

NVIDIA GB10 Grace Blackwell (DGX Spark) · 35B·3B active · NVFP4 · vllm · ctx 65,536

signal score
no private tasks authored yet
reference
0.936
capped at 100 samples · calibration only
single stream
27.6tok/s
TTFT 321ms · what one user feels
peak throughput
321tok/s
at concurrency 32 · what the box serves
921852773691832TOK/S AGGREGATE ↑ / CONCURRENCY →

Throughput scales 11.6× from one stream to 32, but per-stream decode falls to 10.4 tok/s. Batch work and interactive work want opposite settings on this box.

full score card →

The board

latest run per target
targethostsignalreferencetok/s ×1peak tok/smeasured
Qwen3.6-35B-A3B (NVFP4)35B·3B active · NVFP4spark-10.93627.63212026-07-28

Signal scores come from private tasks built from our own workloads. They are the measurement. The set is not published — a public test set gets scraped into the next training run and stops measuring anything.

Reference scores come from public benchmarks, unmodified. They are calibration, not a ranking: if one lands far from its published value, our harness is wrong, not the model.