Skip to content

Elementwise & Reduction

82 ops, 484 workloads — Elementwise 67 · Reduction 15.

One table per op, one row per workload. Ratio is the baseline device time divided by ours in the same measurement regime, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.

Elementwise

VarFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-var
  • W3x: [2048, 4096]dtype=f16hidden-state-var
  • W4x: [64, 32768]dtype=bf16long-seq-var
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W14.49×0.0235torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1052
0.1016
···
W22.39×0.0509torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1230
0.1165
···
W32.47×0.0505torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1232
0.1190
···
W42.33×0.0283torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0671
0.0615
···

SoftplusFwd

  • W1input: [2048, 4096]dtype=bf16mlp-hidden
  • W2input: [2048, 4096]dtype=f16mlp-hidden
  • W3input: [2048, 8192]dtype=bf16mlp-hidden-wide
  • W4input: [2048, 8192]dtype=f16mlp-hidden-wide
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W12.59×0.0601torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1560
0.1383
···
W22.54×0.0602torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5
0.1527
0.1375
0.0276
···
W32.73×0.1067torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2919
0.2740
···
W42.72×0.1057torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5
0.2875
0.2727
0.0510
···

StdFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-std
  • W3x: [2048, 4096]dtype=f16hidden-state-std
  • W4x: [64, 32768]dtype=bf16long-seq-std
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W12.39×0.0253torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0602
0.0960
0.0553
···
W22.69×0.0500torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1333
0.1192
···
W32.60×0.0516torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1343
0.1156
···
W42.58×0.0302torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0790
0.0617
···

HardtanhFwd

  • W1input: [16, 256, 56, 56]dtype=bf16bounded-conv-feat
  • W2input: [16, 256, 56, 56]dtype=f16bounded-conv-feat
  • W3input: [2048, 4096]dtype=bf16bounded-hidden
  • W4input: [2048, 4096]dtype=f16bounded-hidden
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W13.90×0.0510torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1993
0.1953
···
W22.03×0.0530torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5
0.0990
0.0925
0.0419
···
W32.49×0.0473torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.1187
0.1278
···
W41.73×0.0395torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5
0.0669
0.0621
0.0273
···

Relu6Fwd

  • W1input: [1, 4096]dtype=bf16hidden-state-decode
  • W2input: [2048, 4096]dtype=bf16hidden-state-prefill
  • W3input: [2048, 4096]dtype=f16hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.88×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.00375
0.0015
0.002
···
W23.47×0.0428torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active_06 tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.1560
0.0357
0.1374
0.0323
···
W32.00×0.0408torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active_06 tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0848
0.0340
0.0770
0.0300
···

AtanFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W12.25×0.2070torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.4660
0.4592
···
W22.29×0.2025torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.4632
0.4574
···
W32.21×0.2090torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.4622
0.4567
···
W42.32×3.1461torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
7.2988
7.2918
···
W52.34×3.1009torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
7.2707
7.2648
···

Atan2Fwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
  • W7input, other: [4097]dtype=bf16nondiv-tail
  • W8input, other: [4097]dtype=f16nondiv-tail
  • W9input, other: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.43×4.7580torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
2.0466
2.0449
···
W20.45×4.7520torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
2.1224
2.1233
···
W30.37×5.5666torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
2.0465
2.0448
···
W414.03×0.1405torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
1.9691
1.9699
···
W514.27×0.1398torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
1.9939
1.9930
···
W611.27×0.1800torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
2.0281
2.0264
···
W71.94×0.008torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0158
0.0149
0.0152
···
W81.88×0.008torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0150
0.0147
···
W91.65×0.00925torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0150
0.0149
0.0141
···

LerpFwd

  • W1input: [16, 256, 56, 56]end: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]end: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]end: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, end: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, end: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, end: [2048, 4096]dtype=f32hidden-state-prefill
  • W7end, input: [4097]dtype=bf16nondiv-tail
  • W8end, input: [4097]dtype=f16nondiv-tail
  • W9end, input: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W12.61×0.0670torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1810
0.1750
0.1809
···
W23.56×0.0485torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1817
0.1655
0.1696
···
W32.03×0.0865torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1610
0.1588
0.1598
···
W42.07×0.0630torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1240
0.0938
0.1234
···
W52.34×0.0550torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1263
0.0930
0.1171
···
W61.43×0.1022torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1447
0.1220
0.1301
···
W71.89×0.00225torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00225
0.00225
0.00225
···
W82.25×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00225
0.00225
0.00225
···
W91.67×0.00225torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00225
0.00225
0.00225
···

GeluAndMulFwd

x: [M, 28672]

  • W1dtype=bf16M=1ffn-gelu-decode
  • W2dtype=bf16M=2048ffn-gelu-prefill
  • W3dtype=f16M=2048ffn-gelu-prefill

dtype=f16

  • W4x: [17, 514]tail-fp16
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W14.71×0.0035torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0168
0.0123
···
W22.72×0.2345torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.6378
0.8785
···
W31.38×0.2360tilelang-ascend tilelang-ascend basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.3236
0.6368
0.8706
···
W41.05×0.00475ops-nn-ew:aclnnGeluMul open-source basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.0035
0.0160
0.0132
···

GeluTanhAndMulFwd

x: [M, 28672]

  • W1dtype=bf16M=1ffn-gelu-tanh-decode
  • W2dtype=bf16M=2048ffn-gelu-tanh-prefill
  • W3dtype=f16M=2048ffn-gelu-tanh-prefill

dtype=f16

  • W4x: [17, 514]tail-fp16
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W14.48×0.00363torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0163
0.0119
···
W22.75×0.2331torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.6395
0.8731
···
W31.37×0.2355tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.3174
0.6332
0.8704
···
W41.05×0.00475ops-nn-ew:aclnnGeluMul open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.0035
0.0158
0.0132
···

AcosFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W12.08×0.2213torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.4610
0.4521
···
W22.01×0.2198torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.4410
0.4290
···
W31.90×0.2280torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.4323
0.4256
···
W42.11×3.4029torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
7.1878
7.1807
···
W52.03×3.3625torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
6.8243
6.8137
···

MaskedFillFwd

  • W1input: [4096, 4096]mask: [4096, 4096], booldtype=bf16elementwise-16M
  • W2input: [4096, 4096]mask: [4096, 4096], booldtype=f16elementwise-16M
  • W3input: [4096, 4096]mask: [4096, 4096], booldtype=f32elementwise-16M
  • W4input: [16384, 16384]mask: [16384, 16384], booldtype=bf16elementwise-256M
  • W5input: [16384, 16384]mask: [16384, 16384], booldtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.74×0.0798torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1385
0.1249
···
W21.38×0.0819torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1091
0.0945
···
W31.45×0.1338torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1953
0.1800
···
W42.17×1.1293torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
2.4486
2.4352
···
W51.85×1.1447torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
2.1277
2.1126
···

AsinFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.76×0.2170torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.3822
0.3665
···
W21.47×0.2110torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.3098
0.3028
···
W31.31×0.2218torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2898
0.2831
···
W42.18×3.2824torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
7.1515
7.1376
···
W51.48×3.2365torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
4.7953
4.7877
···

CompareFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.86×0.0583torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0960
0.0963
0.0960
···
W22.34×0.0470torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1103
0.1419
0.0983
···
W31.06×0.0794torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0776
0.0780
0.0775
···
W41.81×0.0530torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0973
0.0973
0.0829
···
W51.91×0.0500torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0965
0.0935
0.0815
···
W61.04×0.0922torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0953
0.0925
0.0900
···

HardsigmoidFwd

  • W1input: [32, 240, 1, 1]dtype=bf16mbv3-se-gate
  • W2input: [32, 960, 1, 1]dtype=bf16mbv3-se-gate-deep
  • W3input: [32, 960, 1, 1]dtype=f16mbv3-se-gate-deep
  • W4input: [32, 240, 1, 1]dtype=f16mbv3-se-gate
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.12×0.002torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.00225
0.0015
···
W22.09×0.00275torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.00575
0.00225
···
W32.00×0.00275torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5
0.00575
0.00225
0.0163
···
W41.12×0.002torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5
0.00225
0.0015
0.00487
···

CountNonzeroFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16sparsity-hidden
  • W3x: [2048, 4096]dtype=f16sparsity-hidden
  • W4x: [32, 32768]dtype=f16sparsity-seq
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W14.03×0.0200torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0803
0.0590
···
W21.12×0.1025torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1145
0.0885
···
W30.84×0.1014torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0848
0.0622
···
W41.27×0.0314torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0395
0.0234
···

LogicalAndFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=boolcnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W4input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W5input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=boolhidden-state-prefill
  • W7input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W8input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.85×0.0602torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1025
0.1026
0.1025
···
W21.07×0.0435torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0357
0.0357
0.0357
···
W31.87×0.0525torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0820
0.0820
0.0819
···
W41.47×0.0938torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1324
0.1328
0.1326
···
W51.71×0.0583torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0990
0.0983
0.0841
···
W60.92×0.0376torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0348
0.0318
0.0300
···
W71.46×0.0576torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0830
0.0793
0.0683
···
W81.36×0.1025torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1384
0.1373
0.1227
···

WhereFwd

  • W1condition: [16, 256, 56, 56], boolinput: [256, 1, 1]other: [16, 256, 56, 56]dtype=bf16broadcast
  • W2condition: [16, 256, 56, 56], boolinput: [256, 1, 1]other: [16, 256, 56, 56]dtype=f16broadcast
  • W3condition: [16, 256, 56, 56], boolinput: [256, 1, 1]other: [16, 256, 56, 56]dtype=f32broadcast
  • W15condition: [2, 3, 1], boolinput: [1, 3, 5]other: [2, 1, 5]dtype=bf16three-way-broadcast
  • W16condition: [2, 3, 1], boolinput: [1, 3, 5]other: [2, 1, 5]dtype=f16three-way-broadcast
  • W17condition: [2, 3, 1], boolinput: [1, 3, 5]other: [2, 1, 5]dtype=f32three-way-broadcast
  • W4input: [4096, 4096]dtype=bf16elementwise-16M
  • W5input: [4096, 4096]dtype=f16elementwise-16M
  • W6input: [4096, 4096]dtype=f32elementwise-16M
  • W7input: [16384, 16384]dtype=bf16elementwise-256M
  • W8input: [16384, 16384]dtype=f16elementwise-256M
  • W9condition: [4097], boolinput, other: [4097]dtype=bf16nondiv-tail
  • W10condition: [4097], boolinput, other: [4097]dtype=f16nondiv-tail
  • W11condition: [4097], boolinput, other: [4097]dtype=f32nondiv-tail
  • W12condition: [2048, 4096], boolinput, other: [2048, 4096]dtype=bf16t111-probe-8M
  • W13condition: [2048, 4096], boolinput, other: [2048, 4096]dtype=f16t111-probe-8M
  • W14condition: [2048, 4096], boolinput, other: [2048, 4096]dtype=f32t111-probe-8M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.79×0.0674torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1113
0.1115
0.1110
···
W21.14×0.0668torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0701
0.0703
0.0700
···
W31.21×0.1120torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1253
0.1253
0.1253
···
W41.28×0.1242torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1530
0.1530
0.1527
···
W51.12×0.1227torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1358
0.1242
0.1320
···
W61.10×0.2169torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.2404
0.2405
0.2405
···
W71.25×1.6289torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
2.0532
2.0526
2.0539
···
W81.10×1.6356torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
1.7980
1.7977
1.7984
···
W91.25×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.002
0.002
0.00187
···
W101.00×0.00225torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00175
0.00175
0.00175
···
W111.22×0.00225torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.002
0.002
0.002
···
W121.28×0.0663torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0835
0.0755
0.0793
···
W131.11×0.0654torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0707
0.0660
0.0675
···
W141.09×0.1153torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1253
0.1148
0.1207
···
W153.50×0.0035torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0120
0.0121
0.0120
···
W162.86×0.0035torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00975
0.0095
0.0095
···
W172.35×0.00425torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.00988
0.0100
0.0100
···

AnyFwd

dtype=bool

  • W1x: [4, 128, 4096]dim=[0,2]3d-multidim-reduce
  • W2x: [32, 32768]mask-validation-32k
  • W3x: [32, 4096]mask-validation-4k
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.58×0.0180torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0310
0.0135
···
W21.17×0.0152torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0173
0.00725
···
W31.43×0.0105torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0147
0.00675
···

MishFwd

  • W1input: [16, 256, 80, 80]dtype=bf16yolo-p3
  • W2input: [16, 256, 80, 80]dtype=f16yolo-p3
  • W3input: [16, 512, 40, 40]dtype=bf16yolo-p4
  • W4input: [16, 512, 40, 40]dtype=f16yolo-p4
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.37×0.1966torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:stage4_opt tilelang-mlir-ascend ⓘ eager · n=5
0.2682
0.2607
0.0804
···
W21.37×0.1955torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:stage4_opt tilelang-mlir-ascend ⓘ eager · n=5
0.2667
0.2597
0.0892
···
W31.35×0.1041torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:stage4_opt tilelang-mlir-ascend ⓘ eager · n=5
0.1401
0.1321
0.0470
···
W41.33×0.1047torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:stage4_opt tilelang-mlir-ascend ⓘ eager · n=5
0.1393
0.1315
0.0465
···

LerpTensorFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.37×0.1598torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2165
0.2087
···
W21.35×0.1615torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2160
0.2086
···
W31.15×0.2800torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.3226
0.3190
···
W41.43×2.0451torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
2.9049
2.8920
···
W51.37×2.1231torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
2.9062
2.9065
···

AllFwd

dtype=bool

  • W1x: [4, 128, 4096]dim=[0,2]3d-multidim-reduce
  • W2x: [32, 32768]mask-validation-32k
  • W3x: [32, 4096]mask-validation-4k
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.61×0.0175torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0302
0.0130
···
W21.02×0.0158torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0155
0.00725
···
W31.29×0.0105torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0140
0.0065
···

PowFwd

  • W1input: [16, 256, 56, 56]exponent: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]exponent: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]exponent: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, exponent: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, exponent: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, exponent: [2048, 4096]dtype=f32hidden-state-prefill
  • W7exponent, input: [4097]dtype=bf16nondiv-tail
  • W8exponent, input: [4097]dtype=f16nondiv-tail
  • W9exponent, input: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.76×0.2057torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.3535
0.3540
0.3535
···
W21.65×0.2188torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.3605
0.2831
0.3518
···
W31.78×0.2040torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.3463
0.3461
0.3463
···
W41.49×0.1610torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.2402
0.2325
0.2313
···
W51.09×0.1810tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1857
0.2436
0.2334
0.2310
···
W61.05×0.1358tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1384
0.2489
0.2385
0.2380
···
W71.11×0.0045torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.003
0.003
0.003
···
W80.78×0.00675torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.003
0.003
0.003
···
W91.11×0.0045torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.003
0.003
0.00287
···

CosFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.30×0.2013torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.2607
0.2542
···
W21.24×0.2077tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2640
0.2585
0.2525
···
W31.03×0.2037tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.2025
0.2274
0.2223
···
W41.34×2.9975torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
4.0245
4.0183
···
W51.29×3.1085tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
4.1885
3.9965
3.9905
···

RemainderFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
  • W7input, other: [4097]dtype=bf16nondiv-tail
  • W8input, other: [4097]dtype=f16nondiv-tail
  • W9input, other: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.32×0.1630torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2150
0.2023
···
W21.06×0.1620torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1725
0.1588
···
W31.24×0.1640torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2037
0.1886
···
W41.28×0.1100torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1410
0.1263
···
W51.06×0.1108torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1172
0.1045
···
W61.14×0.1278torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1452
0.1340
···
W71.20×0.00313torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.004
0.00225
···
W81.04×0.00325torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.00325
0.00213
···
W91.15×0.00325torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.002
0.002
···

ClampFwd

  • W1input, min, max: [4096, 4096]dtype=bf16elementwise-16M
  • W2input, min, max: [4096, 4096]dtype=f16elementwise-16M
  • W3input, min, max: [4096, 4096]dtype=f32elementwise-16M
  • W10input, min, max: [16384, 16384]dtype=bf16elementwise-256M
  • W11input, min, max: [16384, 16384]dtype=f16elementwise-256M
  • W4input, max: [4096, 4096]dtype=bf16elementwise-16M-max-only
  • W5input, max: [4096, 4096]dtype=f16elementwise-16M-max-only
  • W6input, max: [4096, 4096]dtype=f32elementwise-16M-max-only
  • W12input, max: [16384, 16384]dtype=bf16elementwise-256M-max-only
  • W13input, max: [16384, 16384]dtype=f16elementwise-256M-max-only
  • W7input, min: [4096, 4096]dtype=bf16elementwise-16M-min-only
  • W8input, min: [4096, 4096]dtype=f16elementwise-16M-min-only
  • W9input, min: [4096, 4096]dtype=f32elementwise-16M-min-only
  • W14input, min: [16384, 16384]dtype=bf16elementwise-256M-min-only
  • W15input, min: [16384, 16384]dtype=f16elementwise-256M-min-only
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.09×0.1575torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1703
0.1656
···
W20.95×0.1591torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1504
0.1491
···
W31.01×0.2770torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.2807
0.2815
···
W41.39×0.1059torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1470
0.1422
···
W51.11×0.1060torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1071
0.1070
···
W61.06×0.2009torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.2100
0.2102
···
W71.53×0.1062torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1489
0.1462
···
W81.10×0.1085torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1185
0.1077
···
W91.07×0.2045torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.2190
0.2115
···
W101.05×2.0339torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
2.1299
2.1202
···
W110.96×2.0513torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
1.9509
1.9507
···
W121.47×1.4291torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
1.8487
1.9780
···
W131.07×1.4126torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
1.5100
1.5101
···
W141.50×1.3942torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
2.1264
2.2654
···
W151.07×1.4265torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
1.5146
1.5129
···

TanhFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.28×0.0818ops-nn-ew-r273:aclnnForeachTanh open-source basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.1030
0.1573
0.1495
···
W21.01×0.0665ops-nn-ew-r273:aclnnForeachTanh open-source basis
tilelang-ascend tilelang-ascend
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.0645
0.2742
0.1557
0.1490
···
W31.06×0.1182ops-nn-ew-r273:aclnnForeachTanh open-source basis
tilelang-ascend tilelang-ascend
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1172
0.2421
0.1475
0.1409
···
W41.28×1.1625ops-nn-ew-r273:aclnnForeachTanh open-source basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
1.4836
2.3478
2.3407
···
W50.96×0.9285ops-nn-ew-r273:aclnnForeachTanh open-source basis
tilelang-ascend tilelang-ascend
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.8892
3.7560
2.3360
2.3301
···

DivFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
  • W7input, other: [4097]dtype=bf16nondiv-tail
  • W8input, other: [4097]dtype=f16nondiv-tail
  • W9input, other: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.00×0.0679torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.0630
0.0630
0.0633
···
W21.49×0.0485torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0726
0.1713
0.0658
···
W31.11×0.0877torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0912
0.0912
0.0912
···
W41.10×0.0578torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0622
0.0587
0.0583
···
W51.07×0.0568tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0765
0.0611
0.0585
0.0583
···
W61.02×0.1020tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0990
0.1037
0.1040
0.1016
···
W71.00×0.00225torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···
W81.12×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···
W91.00×0.00225torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···

SinFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.46×0.1698torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.2487
0.2422
···
W20.90×0.1777ops-nn-ew-r273:aclnnForeachSin open-source basis
tilelang-ascend tilelang-ascend
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1593
0.2239
0.2467
0.2407
···
W30.86×0.1762ops-nn-ew-r273:aclnnForeachSin open-source basis
tilelang-ascend tilelang-ascend
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1457
0.1755
0.2157
0.2102
···
W41.51×2.5425torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
3.8312
3.8260
···
W50.91×2.6194ops-nn-ew-r273:aclnnForeachSin open-source basis
tilelang-ascend tilelang-ascend
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
2.3805
3.5980
3.8037
3.7988
···

SiluAndMulFwd

x: [M, 28672]

  • W1dtype=bf16M=1llama-8b-swiglu-decode
  • W2dtype=bf16M=2048llama-8b-swiglu-prefill
  • W3dtype=f16M=2048llama-8b-swiglu-prefill

dtype=f16

  • W4x: [17, 514]tail-fp16
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.58×0.003ops-nn-ew:aclnnSwiGlu open-source basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.00475
0.0168
0.0115
···
W20.93×0.1994ops-nn-ew:aclnnSwiGlu open-source basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.1802
0.6261
0.8649
···
W30.91×0.2039ops-nn-ew:aclnnSwiGlu open-source basis
tilelang-ascend tilelang-ascend
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.1815
0.3065
0.6272
0.8635
···
W41.00×0.00425ops-nn-ew:aclnnSwiGlu open-source basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.00325
0.0152
0.0130
···

AddFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.12×0.0609torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0679
0.0626
···
W21.18×0.0506torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0587
0.0485
···
W31.10×0.0885torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0983
0.0916
···
W41.10×0.0563torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:example tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0620
0.0578
0.2890
···
W50.96×0.0553tilelang-ascend tilelang-ascend basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0500
0.0540
0.0504
0.1536
···
W60.99×0.1013tilelang-ascend tilelang-ascend basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.0971
0.1035
0.1010
···

SubFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
  • W7input, other: [4097]dtype=bf16nondiv-tail
  • W8input, other: [4097]dtype=f16nondiv-tail
  • W9input, other: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.09×0.0624torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.0615
0.0617
0.0616
···
W21.11×0.0480torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0530
0.0610
0.0471
···
W31.11×0.0872torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0897
0.0897
0.0897
···
W41.09×0.0545torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0590
0.0563
0.0568
···
W51.04×0.0521tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0737
0.0546
0.0542
0.0505
···
W61.03×0.1008tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0985
0.1051
0.1030
0.1010
···
W71.12×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···
W80.88×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···
W91.00×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···

LogicalNotFwd

  • W1input: [4096, 4096]dtype=boolelementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=boolelementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.80×0.0435torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0352
0.0305
···
W21.32×0.0615torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0818
0.0694
···
W31.35×0.0998torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1355
0.1196
···
W40.81×0.5570torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.4525
0.4495
···

MulFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
  • W7input, other: [4097]dtype=bf16nondiv-tail
  • W8input, other: [4097]dtype=f16nondiv-tail
  • W9input, other: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.13×0.0607torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0622
0.0617
0.0620
···
W21.10×0.0498torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0536
0.0517
0.0483
···
W31.08×0.0887torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0892
0.0887
0.0890
···
W41.10×0.0560ops-nn-ew:aclnnForeachMulList open-source ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0583
0.0607
0.0570
0.0536
···
W51.01×0.0540ops-nn-ew:aclnnForeachMulList open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0561
0.0761
0.0548
0.0550
0.0515
···
W61.03×0.1000ops-nn-ew:aclnnForeachMulList open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1032
0.0968
0.1040
0.1017
0.1010
···
W70.89×0.00225ops-nn-ew:aclnnForeachMulList open-source ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.009
0.0015
0.0015
0.0015
···
W80.88×0.002ops-nn-ew:aclnnForeachMulList open-source ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.00887
0.0015
0.0015
0.0015
···
W91.00×0.002ops-nn-ew:aclnnForeachMulList open-source ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.00875
0.0015
0.0015
0.0015
···

Log1pFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.06×0.0675ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0670
0.1060
0.1067
0.0990
···
W21.07×0.0620ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0626
0.1056
0.1062
0.0988
···
W30.99×0.1151ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1151
0.1378
0.1355
0.1310
···
W40.94×0.9560ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.8952
1.5397
1.5394
1.5395
···
W51.02×0.8628ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.8782
1.5441
1.5148
1.5386
···

ReciprocalFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.91×0.0685torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0625
0.0584
···
W20.97×0.0665tilelang-ascend tilelang-ascend ⓘ
torch_npu eager vendor library basis
0.0915
0.0590
···
W30.98×0.1172tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1128
0.1154
0.1115
···
W40.92×0.9778torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.9024
0.9036
···
W51.03×0.8831tilelang-ascend tilelang-ascend ⓘ
torch_npu eager vendor library basis
1.4269
0.8991
···

SignFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.96×0.0755torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0725
0.0643
···
W20.96×0.0700torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0664
0.0578
···
W30.92×0.1260torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1155
0.1092
···
W40.95×1.0490torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.9931
0.9901
···
W50.93×0.9585torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8985
0.8812
···

Expm1Fwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.86×0.0680torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0580
0.0553
···
W20.92×0.0630torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0580
0.0560
···
W30.98×0.1143torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1103
0.1077
···
W40.92×0.9639torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.8835
0.8862
···
W51.04×0.8598torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8889
0.8852
···

SqrtFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.88×0.0707ops-nn-ew:aclnnForeachSqrt open-source
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0669
0.0633
0.0583
···
W20.96×0.0599ops-nn-ew:aclnnForeachSqrt open-source
tilelang-ascend tilelang-ascend
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0622
0.0558
0.0610
0.0535
···
W30.98×0.1144ops-nn-ew:aclnnForeachSqrt open-source
tilelang-ascend tilelang-ascend basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1155
0.1085
0.1116
0.1090
···
W40.93×0.9390ops-nn-ew:aclnnForeachSqrt open-source
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.8890
0.8855
0.8849
···
W50.95×0.8799ops-nn-ew:aclnnForeachSqrt open-source
tilelang-ascend tilelang-ascend basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.8722
0.8336
0.8814
0.8766
···

LogFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.88×0.0700torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0668
0.0576
···
W20.98×0.0639tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.0583
0.0615
0.0587
···
W30.99×0.1153tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1106
0.1128
0.1119
···
W40.92×0.9515torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.8780
0.8821
···
W50.93×0.9524tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.8526
0.8729
0.8776
···

AbsFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.88×0.0668torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0573
0.0540
···
W20.98×0.0602tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0530
0.0553
0.0535
0.1972
···
W30.97×0.1135tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.1060
0.1087
0.1065
···
W40.93×0.9407torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.8788
0.8785
···
W50.94×0.9439tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.8436
0.8655
0.8608
3.4259
···

NegFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.85×0.0705torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0609
0.0565
···
W20.93×0.0610torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0563
0.0535
···
W30.98×0.1150torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1118
0.1098
···
W40.91×0.9720torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.8804
0.8830
···
W51.01×0.8565torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8742
0.8632
···

ExpFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.88×0.0683ops-nn-ew:aclnnForeachExp open-source ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0650
0.0620
0.0597
0.0548
···
W20.94×0.0639ops-nn-ew:aclnnForeachExp open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0631
0.0550
0.0600
0.0624
0.0578
···
W30.97×0.1141ops-nn-ew:aclnnForeachExp open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1145
0.1065
0.1094
0.1105
0.1070
···
W40.93×0.9497ops-nn-ew:aclnnForeachExp open-source ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.8879
0.8811
0.8811
0.8811
···
W50.94×0.8901ops-nn-ew:aclnnForeachExp open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.8778
0.8307
0.8845
0.8781
0.8760
···

Log2Fwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.86×0.0665torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0587
0.0546
···
W20.94×0.0607torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0590
0.0534
···
W30.97×0.1172torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1135
0.1098
···
W40.90×0.9738torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.8884
0.8770
···
W50.98×0.8815torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.8756
0.8696
···

SigmoidFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.82×0.0760torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0620
0.0574
···
W20.98×0.0628tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0885
0.0611
0.0555
0.2163
···
W30.97×0.1150tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1119
0.1111
0.1082
···
W40.85×1.0755torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.9140
0.9096
···
W51.01×0.9009tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
1.4780
0.9135
0.9025
3.7019
···

ReluFwd

  • W1input: [1, 4096]dtype=bf16hidden-state-decode
  • W2input: [2048, 4096]dtype=bf16hidden-state-prefill
  • W3input: [2048, 4096]dtype=f16hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.00×0.00175torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.00175
0.00175
0.00125
0.00325
···
W20.88×0.0395torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0345
0.0338
0.0297
0.0285
···
W30.86×0.0400tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0290
0.0323
0.0350
0.0290
0.0287
···

SiluFwd

  • W1input: [1, 14336]dtype=bf16llama-8b-ffn-decode
  • W2input: [2048, 14336]dtype=bf16llama-8b-ffn-prefill
  • W3input: [2048, 14336]dtype=f16llama-8b-ffn-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.18×0.00275torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.00325
0.0035
0.00175
···
W20.78×0.1406torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.1077
0.1071
0.1030
···
W30.77×0.1385tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.2223
0.1069
0.1062
0.1016
0.3487
···

RsqrtFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.89×0.0710torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0636
0.0583
···
W20.86×0.0705tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.0542
0.0612
0.0574
···
W30.91×0.1225tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1080
0.1128
0.1091
···
W40.90×1.0069torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.9074
0.9107
···
W50.88×1.0249tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.8526
0.9008
0.9019
···

RoundFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.85×0.0715torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0607
0.0565
···
W20.84×0.0700torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0628
0.0553
···
W30.91×0.1236torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1130
0.1085
···
W40.88×1.0065torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.8838
0.8862
···
W50.90×0.9815torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.8799
0.8858
···

TruncFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.83×0.0741torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0622
0.0573
···
W20.86×0.0704torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0602
0.0558
···
W30.90×0.1242torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1118
0.1098
···
W40.88×1.0052torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.8825
0.8858
···
W50.89×0.9839torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.8751
0.8851
···

FloorFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.82×0.0722torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0592
0.0550
···
W20.83×0.0698torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0578
0.0542
···
W30.91×0.1226torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1115
0.1085
···
W40.88×1.0051torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8828
0.8791
···
W50.89×0.9900torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8868
0.8824
···

CeilFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.82×0.0719torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0592
0.0545
···
W20.83×0.0693torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0580
0.0537
···
W30.90×0.1232torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1113
0.1087
···
W40.88×1.0037torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8821
0.8775
···
W50.89×0.9910torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.8888
0.8819
···

HardswishFwd

  • W1input: [32, 96, 56, 56]dtype=bf16mbv3-stage2
  • W2input: [32, 96, 56, 56]dtype=f16mbv3-stage2
  • W3input: [32, 240, 28, 28]dtype=bf16mbv3-stage3
  • W4input: [32, 240, 28, 28]dtype=f16mbv3-stage3
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.89×0.0445torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0395
0.0350
···
W20.84×0.0465torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0398
0.0352
···
W30.81×0.0330torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0272
0.0230
···
W40.82×0.0333torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0275
0.0230
···

BmmFwd

a: [B, M, K]b: [B, K, N]

  • W1dtype=bf16B=64M=128K=2048N=128mha-decode-b64-pv
  • W2dtype=f16B=64M=128K=2048N=128mha-decode-b64-pv
  • W3dtype=bf16B=64M=128K=128N=2048mha-decode-b64-qk
  • W4dtype=f16B=64M=128K=128N=2048mha-decode-b64-qk
  • W5dtype=bf16B=128M=512K=2048N=512moe-prefill-b128
  • W6dtype=bf16B=8M=128K=128N=128small-b8-128
  • W7dtype=f16B=8M=128K=128N=128small-b8-128
  • W8dtype=bf16B=16M=512K=512N=512square-b16-512
  • W9dtype=f16B=16M=512K=512N=512square-b16-512
  • W10dtype=bf16B=32M=256K=256N=256square-b32-256
  • W11dtype=f16B=32M=256K=256N=256square-b32-256
  • W12dtype=bf16B=4M=4096K=4096N=4096square-b4-4k
  • W13dtype=bf16B=8M=1024K=1024N=1024square-b8-1k
  • W14dtype=f16B=8M=1024K=1024N=1024square-b8-1k
  • W15dtype=bf16B=8M=2048K=2048N=2048square-b8-2k
  • W16dtype=f16B=8M=2048K=2048N=2048square-b8-2k
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.78×0.1376catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1010
0.1049
0.1295
0.1045
0.1046
0.1046
···
W20.76×0.1400catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1014
0.1055
0.1313
0.1052
0.1055
0.1052
···
W30.78×0.0926catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0653
0.0675
0.0899
0.0691
0.0690
0.0691
···
W40.76×0.0930catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0650
0.0660
0.0897
0.0668
0.0668
0.0668
···
W50.99×0.7890catlass-r252:r252_bmm open-source ⓘ
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.8089
0.7823
1.5370
0.7802
0.7795
0.7782
···
W60.56×0.0103catlass-r252:r252_bmm open-source ⓘ
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.00675
0.00825
0.00375
0.0120
0.0055
0.0085
···
W70.59×0.0103catlass-r252:r252_bmm open-source ⓘ
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0065
0.00925
0.00375
0.0150
0.0055
0.00925
···
W80.96×0.0500catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0395
0.0488
0.0602
0.0475
0.0475
0.0475
···
W90.90×0.0535catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0420
0.0490
0.0614
0.0490
0.0489
0.0490
···
W100.82×0.0307catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0185
0.0248
0.0227
0.0255
0.0255
0.0255
···
W110.85×0.0293catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0187
0.0243
0.0227
0.0245
0.0246
0.0245
···
W120.96×1.8649catlass-r252:r252_bmm open-source ⓘ
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ basis
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
1.9143
1.7851
5.2805
1.7854
1.7873
1.7857
···
W130.89×0.1126catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0917
0.1101
0.1956
0.1105
0.0925
0.1103
···
W140.87×0.1110catlass-r252:r252_bmm open-source ⓘ basis
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0902
0.1096
0.1953
0.1120
0.0917
0.1095
···
W150.92×0.5693catlass-r252:r252_bmm open-source ⓘ
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ basis
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.5575
0.5215
1.3916
0.5226
0.5230
0.5226
···
W160.92×0.5667catlass-r252:r252_bmm open-source ⓘ
ops-nn-gemm:aclnnBatchMatMul open-source ⓘ
tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.5573
0.5184
1.3915
0.5179
0.5185
0.5185
···

EluFwd

  • W1input: [2048, 4096]dtype=bf16mlp-hidden
  • W2input: [2048, 4096]dtype=f16mlp-hidden
  • W3input: [2048, 8192]dtype=bf16mlp-hidden-wide
  • W4input: [2048, 8192]dtype=f16mlp-hidden-wide
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.78×0.0437torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0340
0.0352
0.0312
0.0332
···
W20.83×0.0419torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0348
0.0352
0.0312
0.0291
···
W30.82×0.0760torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0625
0.0630
0.0583
···
W40.85×0.0752torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0630
0.0649
0.0591
0.0581
···

LeakyReluFwd

  • W1input: [16, 256, 64, 64]dtype=bf16gan-feat
  • W2input: [16, 512, 32, 32]dtype=bf16gan-feat-deep
  • W3input: [16, 512, 32, 32]dtype=f16gan-feat-deep
  • W4input: [16, 256, 64, 64]dtype=f16gan-feat
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.81×0.0710torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0625
0.0550
···
W20.79×0.0420torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0330
0.0288
···
W30.77×0.0415torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0312
0.0285
0.0280
···
W40.82×0.0715torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0587
0.0550
0.0512
···

ErfFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.81×0.1830torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1477
0.1427
···
W20.81×0.1805torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1460
0.1412
···
W30.74×0.1921torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1415
0.1360
···
W40.80×2.7886torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
2.2340
2.2291
···
W50.80×2.7494torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
2.2070
2.2020
···

GeluFwd

  • W1input: [1, 14336]dtype=bf16llama-8b-ffn-decode
  • W2input: [2048, 14336]dtype=bf16llama-8b-ffn-prefill
  • W3input: [2048, 14336]dtype=f16llama-8b-ffn-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.14×0.0035torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00375
0.00325
0.00175
···
W20.58×0.1930torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1123
0.1133
0.1085
···
W30.57×0.1910torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1085
0.1106
0.1050
···

TanFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.51×0.4537torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2295
0.2255
···
W20.52×0.4397torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2280
0.2231
···
W30.46×0.4556torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2115
0.2057
···
W40.50×7.1116torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
3.5642
3.5600
···
W50.52×6.8262torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
3.5307
3.5265
···

CastFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.40×0.1850torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0735
0.0961
···
W20.38×0.1847torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0703
0.0940
···
W30.57×0.2052torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1165
0.1121
···
W40.44×2.8824torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
1.2768
2.5775
···
W50.43×2.8820torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
1.2505
2.5269
···

Im2col

  • W1input: [2, 128, 28, 28]bias: [512]dtype=f16C_out=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1-bias
  • W3input: [2, 512, 28, 28]bias: [128]dtype=f16C_out=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1-bias
  • W5input: [1, 512, 7, 7]bias: [2048]dtype=f16C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1-bias
  • W8input: [1, 256, 112, 112]bias: [512]dtype=f16C_out=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1-bias
  • W10input: [1, 256, 14, 14]bias: [1024]dtype=f16C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1-bias
  • W12input: [1, 64, 56, 56]bias: [128]dtype=f16C_out=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1-bias
  • W16input: [2, 64, 56, 56]bias: [256]dtype=bf16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W17input: [2, 64, 56, 56]bias: [256]dtype=f16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1-bias
  • W20input: [2, 64, 56, 56]bias: [64]dtype=bf16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W21input: [2, 64, 56, 56]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3-bias
  • W24input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2-bias
  • W26input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2-bias
  • W28input: [1, 3, 112, 112]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2-bias
  • W31input: [1, 128, 28, 28]bias: [128]dtype=bf16C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2-bias

input: [N, C_in, H, W]

  • W2dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-expand-1x1
  • W4dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]bottleneck-reduce-1x1
  • W6dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]classifier-1x1
  • W7dtype=f16N=1C_in=2048H=32W=32C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]deeplabv3-aspp-3x3-rate12
  • W9dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]highres-3x3-s1
  • W11dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]late-stage-1x1
  • W13dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]midres-5x5-s1
  • W14dtype=f16N=1C_in=32H=56W=56C_out=32kH=3kW=3stride=[1,1]padding=[1,1]groups=32mobilenetv2-depthwise
  • W15dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W18dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]resnet-1x1
  • W19dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W22dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]resnet-3x3
  • W23dtype=f16N=1C_in=128H=28W=28C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32resnext-grouped-3x3
  • W25dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]stage-transition-3x3-s2
  • W27dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]stage-transition-5x5-s2
  • W29dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]stem-3x3-s2
  • W30dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]stride2
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.02×0.0100torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00625
0.00625
0.00625
···
W21.00×0.0103torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00625
0.00625
0.00625
···
W30.76×0.0235torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0150
0.0110
0.0107
···
W40.76×0.0235torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0150
0.0107
0.0105
···
W50.22×0.0730torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0107
0.0109
0.0107
···
W60.22×0.0732torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0107
0.0107
0.0107
···
W70.02×7.1800torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1583
0.1537
0.1578
···
W80.59×0.3965torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.2442
0.2289
0.2300
···
W90.59×0.3964torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2320
0.2295
···
W100.60×0.0175torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00637
0.00625
0.00637
···
W110.60×0.0174torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0065
0.00637
0.00625
···
W121.19×0.1928torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.2300
0.2281
0.2274
···
W131.13×0.1948torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.2214
0.2255
···
W141.99×0.0410torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0739
0.0686
0.0741
···
W153.13×0.00975torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0219
0.0227
0.0225
···
W163.30×0.0100torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0249
0.0245
0.0243
···
W173.04×0.0104torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0240
0.0249
0.0240
···
W183.05×0.0103torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0236
0.0240
0.0231
···
W191.49×0.1118torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1594
0.1479
···
W201.50×0.1110torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1678
0.1370
0.1466
···
W211.50×0.1103torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1725
0.1704
0.1608
···
W221.50×0.1100torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1645
0.1555
···
W230.76×0.0675torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0467
0.0474
0.0460
···
W240.02×2.8989torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0616
0.0442
0.0449
···
W250.02×2.8984torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0534
0.0437
···
W260.01×8.2525torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1316
0.1133
0.1140
···
W270.01×8.2511torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1242
0.1156
···
W280.24×0.2868torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0678
0.0670
0.0727
···
W290.24×0.2864torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.0678
0.0781
···
W300.01×0.7562torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0105
0.00775
0.0075
···
W310.01×0.7576torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0105
0.00725
0.00725
···

BiasAddFwd

  • W1input: [16, 256, 56, 56]bias: [56]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]bias: [56]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]bias: [56]dtype=f32cnn-feat-broadcast
  • W4input: [2048, 4096]bias: [4096]dtype=bf16hidden-state-prefill
  • W5input: [2048, 4096]bias: [4096]dtype=f16hidden-state-prefill
  • W6input: [2048, 4096]bias: [4096]dtype=f32hidden-state-prefill
  • W7bias, input: [4097]dtype=bf16nondiv-tail
  • W8bias, input: [4097]dtype=f16nondiv-tail
  • W9bias, input: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.01×8.1152torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0545
0.0545
0.0542
···
W20.01×8.1036torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0550
0.0490
0.0481
···
W30.09×1.0171torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0886
0.0885
0.0887
···
W41.03×0.0390torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0393
0.0362
0.0350
···
W50.96×0.0385torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0367
0.0350
0.0305
···
W60.98×0.0650torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0631
0.0610
0.0585
···
W71.00×0.00225torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0015
0.0015
0.0015
···
W80.88×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0015
0.00125
0.0015
···
W91.00×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0015
0.0015
0.0015
···

ProdFwd

  • W1x: [2048, 4096]dtype=bf16hidden-state-reduce
  • W2x: [2048, 4096]dtype=f16hidden-state-reduce
  • W3x: [64, 32768]dtype=bf16long-seq-reduce
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.20×0.4052torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0779
0.0650
···
W20.20×0.4060torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0789
0.0636
···
W3—·torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0376
0.0194
···

GemvFwd

a: [M, K]x: [K]

trans_a=false

  • W1dtype=bf16M=128K=2048n=7168trans_b=trueds-v3-decode-down
  • W2dtype=bf16M=128K=7168n=2112trans_b=trueds-v3-decode-gate-up
  • W3dtype=bf16M=4096K=7168n=4096trans_b=trueds-v3-prefill-attn-proj
  • W4dtype=f16M=4096K=7168n=4096trans_b=trueds-v3-prefill-attn-proj
  • W5dtype=bf16M=4096K=2048n=7168trans_b=trueds-v3-prefill-down
  • W6dtype=bf16M=4096K=7168n=2112trans_b=trueds-v3-prefill-gate-up
  • W7dtype=bf16M=4096K=16384n=7168trans_b=truek-dominant-7168x16384
  • W8dtype=bf16M=16K=7168n=4096trans_b=truemid-m16-attn
  • W9dtype=bf16M=32K=7168n=4096trans_b=truemid-m32-attn
  • W10dtype=bf16M=64K=2048n=7168trans_b=truemid-m64-down
  • W11dtype=bf16M=96K=7168n=2112trans_b=truemid-m96-gate-up
  • W12dtype=bf16M=1024K=1024n=1024trans_b=falsesquare-1k-nn
  • W13dtype=f16M=1024K=1024n=1024trans_b=falsesquare-1k-nn
  • W14dtype=bf16M=4096K=1536n=24576trans_b=truewide-n-24576
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.41×0.0777torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0320
0.0213
···
W20.14×0.2875torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0403
0.0269
···
W30.08×2.0386torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2107
0.1679
···
W40.09×1.8646torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1670
0.1635
···
W50.09×1.0717torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0930
0.0895
···
W60.09×1.8217torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1708
0.1689
···
W70.07×3.0372torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.2214
0.2199
0.2201
···
W80.10×0.2890torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0248
0.0245
0.0249
···
W90.14×0.2772torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0250
0.0245
0.0245
···
W100.32×0.0943torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0198
0.0195
0.0196
···
W110.11×0.3619torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0261
0.0262
0.0261
···
W120.13×0.2451torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0325
0.0238
···
W130.12×0.2387torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0377
0.0235
···
W140.09×0.8599torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0750
0.0750
0.0750
···

MedianFwd

  • W1x: [2048, 4096]dtype=bf16hidden-state-reduce
  • W2x: [2048, 4096]dtype=f16hidden-state-reduce
  • W3x: [64, 32768]dtype=bf16long-seq-reduce
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.00×363.6075torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.5075
0.5055
···
W20.00×363.6019torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.5104
0.5042
···
W30.00×117.0786torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2700
0.2632
···

SortFwd

  • W1x: [2048, 4096]dtype=bf16hidden-state-reduce
  • W2x: [2048, 4096]dtype=f16hidden-state-reduce
  • W3x: [64, 32768]dtype=bf16long-seq-reduce
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.00×363.5121torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.3936
0.3886
···
W20.00×363.5584torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.3915
0.3937
···
W30.00×117.0429torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.2110
0.2037
···

DropoutFwd

Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by

Reduction

CummaxFwd

  • W1x: [2048, 4096]dtype=bf16hidden-state-scan
  • W2x: [2048, 4096]dtype=f16hidden-state-scan
  • W3x: [64, 32768]dtype=bf16long-seq-scan
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W15.13×13.8046torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
70.8431
70.8659
···
W25.13×13.8078torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
70.8836
70.8146
···
W33.52×5.1425torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
18.1371
17.7641
···

MeanVarWelfordFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-var-mean
  • W3x: [2048, 4096]dtype=f16hidden-state-var-mean
  • W4x: [64, 32768]dtype=bf16long-seq-var-mean
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W14.44×0.0260torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1161
0.1087
···
W22.77×0.0530torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1437
0.1391
···
W32.82×0.0530torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1447
0.1385
···
W43.07×0.0271torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0798
0.0740
···

LogSumExpFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [32, 32, 32768]dtype=bf16attn-weights-32k
  • W3x: [32, 32, 4096]dtype=bf16attn-weights-4k
  • W4x: [32, 32, 4096]dtype=f16attn-weights-4k
  • W5x: [4, 102400]dtype=bf16lm-head-logits
  • W6x: [4, 102400]dtype=f16lm-head-logits
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.70×0.0423torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5
0.0721
0.0660
0.0235
···
W22.75×0.1708torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5
0.4632
0.4584
0.1111
···
W32.77×0.0357torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5
0.1065
0.0975
0.0183
···
W42.79×0.0352torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5
0.0980
0.0960
0.0168
···
W52.79×0.0205torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5
0.0573
0.0515
0.0139
···
W63.56×0.0158torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5
0.0560
0.0500
0.0144
···

SumFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-reduce
  • W3x: [2048, 4096]dtype=bf16dim=0hidden-state-reduce-dim0
  • W4x: [2048, 4096]dtype=f16hidden-state-reduce
  • W5x: [2048, 4096]dtype=bf16dim=-1keepdim=truehidden-state-reduce-keepdim
  • W6x: [64, 32768]dtype=bf16long-seq-reduce
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.94×0.0350torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0678
0.0520
···
W21.85×0.0372torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0688
0.0527
···
W31.08×0.0870torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0932
0.0831
···
W41.79×0.0375torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0686
0.0515
···
W52.17×0.0348torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0755
0.0640
···
W61.70×0.0205torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0344
0.0200
···

AminFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-reduce
  • W3x: [2048, 4096]dtype=f16hidden-state-reduce
  • W4x: [64, 32768]dtype=bf16long-seq-reduce
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.88×0.0357torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0335
0.0217
···
W21.84×0.0372torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0651
0.0542
···
W31.82×0.0375torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0670
0.0542
···
W41.37×0.0230torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0333
0.0205
···

MeanFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-reduce
  • W3x: [2048, 4096]dtype=f16hidden-state-reduce
  • W4x: [64, 32768]dtype=bf16long-seq-reduce
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.77×0.0365torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0280
0.0565
0.0208
···
W21.65×0.0390torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0649
0.0600
0.0535
···
W31.76×0.0374torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0660
0.0607
0.0535
···
W41.72×0.0198torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0340
0.0312
0.0226
···

AmaxFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-reduce
  • W3x: [2048, 4096]dtype=f16hidden-state-reduce
  • W4x: [64, 32768]dtype=bf16long-seq-reduce
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.80×0.0357torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0290
0.0215
···
W21.85×0.0357torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0655
0.0515
···
W31.81×0.0365torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0649
0.0515
···
W41.41×0.0220torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0330
0.0192
···

MaximumFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
  • W7input, other: [4097]dtype=bf16nondiv-tail
  • W8input, other: [4097]dtype=f16nondiv-tail
  • W9input, other: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.23×0.0626torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0651
0.0653
0.0650
···
W21.06×0.0498torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0530
0.0678
0.0485
···
W31.14×0.0890torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0899
0.0897
0.0899
···
W41.09×0.0585torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0643
0.0580
0.0592
···
W50.98×0.0573tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0851
0.0558
0.0565
0.0519
···
W61.02×0.1026tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.1000
0.1067
0.1045
0.1022
···
W71.12×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···
W80.88×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.00125
0.00125
0.00125
···
W91.12×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···

MinimumFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
  • W7input, other: [4097]dtype=bf16nondiv-tail
  • W8input, other: [4097]dtype=f16nondiv-tail
  • W9input, other: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.28×0.0628torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0633
0.0630
0.0633
···
W21.02×0.0515torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0519
0.0575
0.0485
···
W31.12×0.0870torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0890
0.0891
0.0890
···
W41.08×0.0560torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0610
0.0566
0.0580
···
W50.99×0.0548tilelang-ascend tilelang-ascend ⓘ
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0849
0.0545
0.0544
0.0510
···
W61.03×0.1003tilelang-ascend tilelang-ascend ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0985
0.1047
0.1032
0.1016
···
W70.89×0.00225torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···
W80.88×0.002torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.00125
0.0015
···
W91.00×0.00225torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0015
0.0015
0.0015
···

ArgmaxFwd

  • W1x: [4, 128, 4096]dtype=f16dim=03d-non-last-axis-argmax
  • W2x: [2048, 4096]dtype=bf16dim=-1hidden-state-argmax
  • W3x: [2048, 4096]dtype=f16dim=-1hidden-state-argmax
  • W4x: [4, 102400]dtype=bf16dim=-1lm-head-argmax
  • W5x: [4, 102400]dtype=f16dim=-1lm-head-argmax
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.08×0.0295torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0232
0.0232
0.0231
···
W20.90×0.0521torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0465
0.0460
0.0466
···
W30.75×0.0522torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0390
0.0380
0.0387
···
W41.01×0.0174torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0175
0.0213
0.0173
···
W50.59×0.0175torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0103
0.0145
0.00975
···

ArgminFwd

dim=-1

  • W1x: [2048, 4096]dtype=bf16hidden-state-argmin
  • W2x: [2048, 4096]dtype=f16hidden-state-argmin
  • W3x: [4, 102400]dtype=bf16lm-head-argmin
  • W4x: [4, 102400]dtype=f16lm-head-argmin
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.93×0.0542torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0503
0.0460
0.0454
···
W20.70×0.0540torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0385
0.0376
0.0367
···
W30.86×0.0203torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0175
0.0213
0.0163
···
W40.49×0.0187torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.00925
0.0138
0.0085
···

SegmentSumFwd

  • W1x: [2048, 4096]dtype=bf16hidden-state-reduce
  • W2x: [2048, 4096]dtype=f16hidden-state-reduce
  • W3x: [64, 32768]dtype=bf16long-seq-reduce
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.27×0.2995torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0809
0.0775
···
W20.28×0.3003torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0875
0.0780
···
W30.47×0.0975torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0457
0.0400
···

LogSoftmaxFwd

  • W1x: [32, 32, 32768]dtype=bf16attn-weights-32k
  • W2x: [32, 32, 4096]dtype=bf16attn-weights-4k
  • W3x: [32, 32, 4096]dtype=f16attn-weights-4k
  • W4x: [32, 32, 4096]dtype=f32attn-weights-4k
  • W5x: [4, 102400]dtype=bf16lm-head-logits
  • W6x: [4, 102400]dtype=f16lm-head-logits
  • W7x: [4, 102400]dtype=f32lm-head-logits
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.63×0.8110torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.4966
0.4846
0.4911
···
W20.65×0.1275torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0850
0.0665
0.0664
···
W30.67×0.1231torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0815
0.0668
0.0660
···
W40.36×0.1618torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0583
0.0540
0.0460
···
W50.16×0.3415torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0408
0.0406
0.0408
···
W60.16×0.3337torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0405
0.0408
0.0405
···
W70.10×0.3897torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0374
0.0375
0.0372
···

SoftmaxFwd

  • W1x: [32, 32, 32768]dtype=bf16attn-weights-32k
  • W2x: [32, 32, 4096]dtype=bf16attn-weights-4k
  • W3x: [32, 32, 4096]dtype=f16attn-weights-4k
  • W4x: [32, 32, 4096]dtype=f32attn-weights-4k
  • W5x: [4, 102400]dtype=bf16lm-head-logits
  • W6x: [4, 102400]dtype=f16lm-head-logits
  • W7x: [4, 102400]dtype=f32lm-head-logits
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.57×0.8177torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.4669
0.4818
0.4929
···
W20.55×0.1283torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0727
0.0630
0.0658
···
W30.55×0.1269torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0695
0.0641
0.0653
···
W40.31×0.1613torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0495
0.0532
0.0474
···
W50.15×0.3559torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0425
0.0423
0.0423
···
W60.15×0.3548torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0423
0.0423
0.0419
···
W70.09×0.4040torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
tilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0380
0.0381
0.0382
0.1042
···

MaskedReduceSumFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-reduce
  • W3x: [2048, 4096]dtype=bf16dim=0hidden-state-reduce-dim0
  • W4x: [2048, 4096]dtype=f16hidden-state-reduce
  • W5x: [2048, 4096]dtype=bf16dim=-1keepdim=truehidden-state-reduce-keepdim
  • W6x: [64, 32768]dtype=bf16long-seq-reduce
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.16×0.6412torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0995
0.0848
···
W20.07×2.3206torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1578
0.1417
···
W30.08×2.3537torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1747
0.1653
···
W40.07×2.3205torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1562
0.1405
···
W50.08×2.3161torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1570
0.1469
···
W60.11×0.6468torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0737
0.0624
···