Does this model earn
its place on our hardware?
Two numbers decide a self-hosting call: whether a model is good enough at the work we actually do, and whether it is fast enough on the box we would run it on. Most leaderboards report only the first. Every card here reports both, measured together, on the same machine.
Qwen3.6-35B-A3B (NVFP4)
NVIDIA GB10 Grace Blackwell (DGX Spark) · 35B·3B active · NVFP4 · vllm · ctx 65,536
Throughput scales 11.6× from one stream to 32, but per-stream decode falls to 10.4 tok/s. Batch work and interactive work want opposite settings on this box.
full score card →The board
latest run per target| target | host | signal | reference | tok/s ×1 | peak tok/s | measured |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B (NVFP4)35B·3B active · NVFP4 | spark-1 | — | 0.936 | 27.6 | 321 | 2026-07-28 |
Signal scores come from private tasks built from our own workloads. They are the measurement. The set is not published — a public test set gets scraped into the next training run and stops measuring anything.
Reference scores come from public benchmarks, unmodified. They are calibration, not a ranking: if one lands far from its published value, our harness is wrong, not the model.