Skip to content

Benchmarks

Nightly snapshot

GPU Ascend 910B1 · commit 250b2be6447e · run date 2026-09-23T09:32:42+00:00 · 149 ops, 1159 workloads · nightly run

Page rendered 2026-10-08 01:43 UTC from the latest snapshot.

The first phase is the acceptance list now

Since 2026-09-23 this overview and the per-family data pages are scoped to the 149 operators of oplist150, the first-phase acceptance list (docs/reports/R267-data/oplist150.json, 150 rows; an operator that fills two rows is counted once) — the list progress is reported against. Until then these pages were scoped to the 158 operators of docs/tasks/OPS-178.md that D018 kept. What changed in these counts on that day is the scope, not performance — neither a regression nor a gain: no ratio was re-measured for the change, and moving an operator from one page to another changes none of its numbers.

Operators outside the list, Attention among them, are on the second-phase page when the snapshot publishes a ratio for them.

Environment

timer torch_npu_profiler_device_kernel_interval_union
platform Linux-5.10.0-136.12.0.86.r1526_92.hce2.aarch64-aarch64-with-glibc2.34
python 3.9.9
provider tilelang
headline_regime graph
harness_commit c04ff24867ce4b8e59ef426e67e0518d807b9bff
backend_commit 250b2be6447e8c2699940a176103929956b99daa

Not published by this run: image, driver, cuda, torch, tilelang.

Method

  • One process, common inputs. Every implementation of an op is timed on the same tensors in the same process, in forward and then reversed order so drift does not land on whichever ran last.
  • Ascend nightly sampling: WARMUP=10 / REPEATS=30 per implementation, reporting the median of the measured samples, with the 192 MiB L2 evicted between iterations.
  • Compilation and workspace setup excluded.
  • Device time is what is compared — the union of the intervals the device spent executing the call's kernels, collected through torch_npu.profiler. A run that cannot collect device activity fails rather than falling back to a different clock.
  • MLIR opponents are measured in a subprocess and corrected by the ratio of common case-yardstick medians, with fresh case and fixed global yardsticks in every round. They enter the eager pool only. Each candidate shows its actual N; N<5 or unknown counts remain provisional and do not count as attainment.

Coverage

  • 146 of 149 ops are rated against a real alternative measured on the identical workload. The denominator is the displayed target set; unconnected targets retain empty tables. Unrated rows have no measured comparison or only an eager reference.
  • Each op is raced against a pool. The fastest eligible candidate in the headline measurement regime becomes its baseline. The row lists all candidates with their individual times and provenance tiers.
  • 20 of 149 ops have a tier-1 open-source baseline in their pool at all — a kernel built from the source of a third-party Ascend operator library, admitted only on evidence that the library's own compiled kernel ran. Same denominator as above. This is the stricter of the two readings, and the one to quote for a claim about open-source libraries.
  • 24/149 ops have a measured tilelang-ascend candidate (tier tilelang_ref) in the published pool; this count can overlap the open-source count above.
  • For the other 102 of 149, the pool holds vendor implementations only — a CANN built-in, or torch_npu's own dispatch. Those rows are still measured against a real implementation of the op on the identical workload, but a win there is not a win over an open-source library. The tier badge on each line says which it was.
  • Absent from every table: 0 workloads errored and 3 were skipped in this run.

How these numbers are taken

Data

Page Ops Workloads
Linear Attention & SSM 2 10
GEMM, MoE & Quantization 20 251
Elementwise & Reduction 82 484
Norm, Conv, Pool & Other 45 414