Skip to content

How these numbers are taken

Every data page answers one question: how does TileOPs compare to the fastest alternative implementation of the same op, on the same workload? Each op gets one table, with one row per workload. Nothing is averaged across workloads: every number on the page belongs to a single shape and dtype.

Who the comparison is against

Ratio = baseline device time ÷ our device time, for the same measurement regime and each (op, workload, dtype). Both sides are selected for that row’s own shape and dtype.

The baseline document decides which implementations may enter the comparison group. Measurement selects the fastest eligible candidate in the headline regime on the exact workload. The Alternatives column lists the candidates and marks the selected baseline as basis.

The Alternatives column lists every candidate's name, provenance tier and measured device time, including open-source candidates. Each row displays one ratio in the headline measurement regime. MLIR candidates marked eager were measured without graph capture; their times do not replace a captured graph baseline.

A pool of one is written single_candidate: only one opponent could be built for that workload. It is still a measured comparison — there was simply no race to run.

What a tier means

Every line in Alternatives carries a tier. A tier records where that kernel came from. It says nothing about which kernel is faster.

Tier Meaning
open-source A kernel out of an open-source Ascend library, reached through that library's own entry point. Admitted only on evidence that the library's own compiled kernel actually ran: a custom operator package that is missing a kernel falls back to the CANN built-in silently, with the call succeeding, the outputs correct and the timings plausible. So provenance here is established by tracing which binary the process opened, never by the call returning.
tilelang-ascend Original example or test kernel from tilelang-ascend (Tile-AI), compiled with the same DSL/compiler. Admitted only after the harness source-provenance and numeric gates. This tier describes kernel origin, not speed.
vendor library A vendor implementation: a CANN built-in operator, or torch_npu's own dispatch for the op. A real implementation on the identical workload, and on this device frequently the fastest one.
torch_compile The same PyTorch program handed to torch.compile on this backend (aclgraph or ge), timed as one more candidate in the pool. Not a hand-written kernel and not a vendor dispatch: a compiler's answer to the same problem. It carries the neutral badge because it is neither of the two, and it wins often enough here that a reader should know what beat us when it does.
agent-written kernel A kernel an agent wrote, admitted to the pool by D060 under the same source-provenance and numeric gates as the others. Like every tier here it describes where the kernel came from, not how fast it is. No row in the current snapshot carries it.

So open-source beside a slower time than vendor library on the same row is not an error, and for several op families here it is the ordinary case. Reading the badge as a strength ranking is the one mistake this column exists to prevent.

The historical inventory covers 91 elementwise, reduction, scan and dropout ops. Counted on the snapshot this page is rendered from: 32/91 have at least one timed source candidate (handwritten or tilelang_ref), 58/91 have none, and 1/91 cannot be determined because this run publishes no workload for them. These are measured pool observations after the shape/dtype contract gate, not an exhaustive source inventory or a numerical-correctness pass rate. Refusals and workload coverage remain case-specific; beating a vendor implementation alone does not establish a win over an open-source kernel.

Which regime a row was measured in

A ratio is comparable with another ratio only if the two were taken the same way. Every workload here is measured in a named regime, and this run's headline regime is graph — the regime the coverage bullets, the colours and the op orderings are all about.

A few operators have no measurement in that regime and are published from an older record that only ever measured eager. Those rows are kept: the number is a real measurement of a real workload, and dropping them would shorten the page with no explanation. They are named instead.

eager — this row's ratio was taken in eager, not in graph. Read it against other eager rows, never against the ones around it.

The same rule holds for every count on the overview: ops measured in graph and ops measured otherwise are reported as two numbers and are never added into one.

The colour is the verdict

Meaning
0.74× Slower than the alternative — below 0.95×.
1.02× Level with it — 0.95–1.05×, inside measurement noise.
1.42× Faster than it — 1.05× and above.
18.06× No fast alternative to compare against — only a functional reference (a name ending in -ref). The number means little.
— No alternative at all ran on this workload.

A ratio is the alternative's device time divided by ours, so above 1 means TileOPs is faster.

Columns

Column Meaning
Workload W1, W2, … — the key above each table spells each one out: the benchmark's own id for it, the dtype it ran at, and every input tensor as name: shape, dtype. Tensors sharing a shape are named together, and each carries its own dtype, so a mask in bool says so where it is read. After the tensors come the dimensions the op is sized by rather than shaped by (m, n, k for a GEMM, num_experts for MoE routing), then dimmed, the parameters the call did not leave at the signature's default. A quantity the others already fix — max_seqlen_q is max(q_lens) — is not repeated.
Ratio alt / ours — the baseline's device time divided by ours, the one number the colour grades. Candidate names, tiers and individual times are listed in Alternatives.
Device time Milliseconds the device spent executing the call's kernels — the union of their intervals. Every comparison on these pages uses it.
Alternatives One line per candidate, with its device time in ms and provenance tier. The line marked basis supplies the baseline for the headline ratio. Hover a shortened name to see its full binding. Times from different measurement regimes, including candidates marked eager, must be read in their own regime. Where a run publishes no pool, the lines are the named baselines it timed: tuned library kernels (fla, mamba, fa3, triton, …), native PyTorch ops (torch), or eager reference compositions whose names end in -ref.
Throughput TFLOP/s: required FLOPs / device time. The count is analytic — the op's eval_roofline formula evaluated on the workload's own shapes, not a hardware counter — so it counts the work the problem demands, not the instructions the kernel issued. Padding, recompute or a masked-out tile is therefore invisible here, and the figure is only comparable between implementations of the same op on the same workload.
SOL Share of the algorithmic speed-of-light: the fastest time physics allows for the workload, divided by our device time. The Ratio column says whether someone is faster today; SOL says how much faster anyone could ever be. Details below.
Bound The resource that sets the workload's floor: mem (HBM traffic), comp (compute throughput), or lat — the workload is too small for the model to judge, and its SOL number greys out with it.

How that device time is measured — what it counts, what it leaves out, and where it refuses to produce a number — is in Benchmark Timing.

Each op's heading carries its workload count and its test outcome (✅ passed · ❌ failed · ⏭️ all skipped · · no test matched).

Speed of light

SOL is algorithmic speed-of-light efficiency: max(bytes / bandwidth, FLOPs / compute roof) / device time, priced against the machine's calibrated ceilings — the bandwidth and compute rates microbenchmarks actually reach on this GPU, not the spec sheet. 100% means no implementation of this algorithm on this hardware can be faster.

Three statements delimit what a reading means:

  1. Bytes are the algorithm's minimum traffic — each input read once, each output written once — not the DRAM traffic the kernel generated. A kernel that moves data twice scores low; that is the point.
  2. FLOPs follow the TileOPs counting convention (a transcendental counts as one), not per-instruction hardware cost; the metric does not certify a special-function-bound kernel as at its limit.
  3. The compute roof is the unit an optimal implementation would use — declared per op, never inferred from the running kernel, so a kernel on the wrong unit is measured against the right ceiling.
SOL Bound Meaning
92% mem At the achievable ceiling (from 80%). The ceiling is an envelope over access mixes, and a kernel's own mix caps below it, so the line leaves room for every mix. Optimizing further buys at most the remainder.
63% mem Headroom remains.
41% lat The workload is too small for the model to judge — launch overhead dominates the measurement, not the roofline — so the number is shown but not graded.
⚠ 108% mem Above the calibrated ceiling: the formula or the calibration is wrong. Never read it as a fast kernel.
· · An input is missing: no roofline formula, a timing not collected on the device, or no GPU profile for the device.

The model, its thresholds and the formula-audit machinery are specified in TileOPs docs/design/roofline.md; the page imports that implementation rather than re-deriving it.

Where the shapes come from

The shapes are read from the TileOPs spec manifest, joined to a row by the label and dtype the benchmark id is built from — so a row states what its op is declared to take. A workload the manifest does not declare — an edge-case probe written by hand rather than driven by a spec — instead shows the shape the snapshot recorded for that run, which is one measured call rather than a declaration. Where neither is available, the row shows its benchmark id alone, with no shapes under it.

Empty cells

· means an input to that metric was not recorded, never that the value is zero: the op reported no FLOP count for that workload, or no alternative ran on it.

In the comparison group the same absence has its own spelling. A candidate the run selected but never measured is recorded as not-timed. Where the run published that baseline's time elsewhere, the cell shows it followed by *, with the reason on the cell; where it did not, the cell is ·. Neither is ever rendered as 0, which would read as an infinitely fast kernel.