Elementwise & Reduction¶
82 ops, 484 workloads — Elementwise 67 · Reduction 15.
One table per op, one row per workload. Ratio is the baseline device time divided by ours in the same measurement regime, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.
Elementwise¶
VarFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-var - W3x: [2048, 4096]dtype=f16
hidden-state-var - W4x: [64, 32768]dtype=bf16
long-seq-var
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.49× | 0.0235 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1052 0.1016 | · | · | · |
| W2 | 2.39× | 0.0509 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1230 0.1165 | · | · | · |
| W3 | 2.47× | 0.0505 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1232 0.1190 | · | · | · |
| W4 | 2.33× | 0.0283 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0671 0.0615 | · | · | · |
SoftplusFwd¶
- W1input: [2048, 4096]dtype=bf16
mlp-hidden - W2input: [2048, 4096]dtype=f16
mlp-hidden - W3input: [2048, 8192]dtype=bf16
mlp-hidden-wide - W4input: [2048, 8192]dtype=f16
mlp-hidden-wide
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.59× | 0.0601 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1560 0.1383 | · | · | · |
| W2 | 2.54× | 0.0602 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 | 0.1527 0.1375 0.0276 | · | · | · |
| W3 | 2.73× | 0.1067 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2919 0.2740 | · | · | · |
| W4 | 2.72× | 0.1057 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 | 0.2875 0.2727 0.0510 | · | · | · |
StdFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-std - W3x: [2048, 4096]dtype=f16
hidden-state-std - W4x: [64, 32768]dtype=bf16
long-seq-std
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.39× | 0.0253 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0602 0.0960 0.0553 | · | · | · |
| W2 | 2.69× | 0.0500 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1333 0.1192 | · | · | · |
| W3 | 2.60× | 0.0516 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1343 0.1156 | · | · | · |
| W4 | 2.58× | 0.0302 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0790 0.0617 | · | · | · |
HardtanhFwd¶
- W1input: [16, 256, 56, 56]dtype=bf16
bounded-conv-feat - W2input: [16, 256, 56, 56]dtype=f16
bounded-conv-feat - W3input: [2048, 4096]dtype=bf16
bounded-hidden - W4input: [2048, 4096]dtype=f16
bounded-hidden
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 3.90× | 0.0510 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1993 0.1953 | · | · | · |
| W2 | 2.03× | 0.0530 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 | 0.0990 0.0925 0.0419 | · | · | · |
| W3 | 2.49× | 0.0473 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.1187 0.1278 | · | · | · |
| W4 | 1.73× | 0.0395 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 | 0.0669 0.0621 0.0273 | · | · | · |
Relu6Fwd¶
- W1input: [1, 4096]dtype=bf16
hidden-state-decode - W2input: [2048, 4096]dtype=bf16
hidden-state-prefill - W3input: [2048, 4096]dtype=f16
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.88× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.00375 0.0015 0.002 | · | · | · |
| W2 | 3.47× | 0.0428 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active_06 tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.1560 0.0357 0.1374 0.0323 | · | · | · |
| W3 | 2.00× | 0.0408 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active_06 tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0848 0.0340 0.0770 0.0300 | · | · | · |
AtanFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.25× | 0.2070 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.4660 0.4592 | · | · | · |
| W2 | 2.29× | 0.2025 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.4632 0.4574 | · | · | · |
| W3 | 2.21× | 0.2090 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.4622 0.4567 | · | · | · |
| W4 | 2.32× | 3.1461 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 7.2988 7.2918 | · | · | · |
| W5 | 2.34× | 3.1009 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 7.2707 7.2648 | · | · | · |
Atan2Fwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill - W7input, other: [4097]dtype=bf16
nondiv-tail - W8input, other: [4097]dtype=f16
nondiv-tail - W9input, other: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.43× | 4.7580 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 2.0466 2.0449 | · | · | · |
| W2 | 0.45× | 4.7520 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 2.1224 2.1233 | · | · | · |
| W3 | 0.37× | 5.5666 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 2.0465 2.0448 | · | · | · |
| W4 | 14.03× | 0.1405 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 1.9691 1.9699 | · | · | · |
| W5 | 14.27× | 0.1398 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 1.9939 1.9930 | · | · | · |
| W6 | 11.27× | 0.1800 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 2.0281 2.0264 | · | · | · |
| W7 | 1.94× | 0.008 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0158 0.0149 0.0152 | · | · | · |
| W8 | 1.88× | 0.008 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0150 0.0147 | · | · | · |
| W9 | 1.65× | 0.00925 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0150 0.0149 0.0141 | · | · | · |
LerpFwd¶
- W1input: [16, 256, 56, 56]end: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]end: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]end: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, end: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, end: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, end: [2048, 4096]dtype=f32
hidden-state-prefill
- W7end, input: [4097]dtype=bf16
nondiv-tail - W8end, input: [4097]dtype=f16
nondiv-tail - W9end, input: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.61× | 0.0670 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1810 0.1750 0.1809 | · | · | · |
| W2 | 3.56× | 0.0485 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1817 0.1655 0.1696 | · | · | · |
| W3 | 2.03× | 0.0865 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1610 0.1588 0.1598 | · | · | · |
| W4 | 2.07× | 0.0630 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1240 0.0938 0.1234 | · | · | · |
| W5 | 2.34× | 0.0550 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1263 0.0930 0.1171 | · | · | · |
| W6 | 1.43× | 0.1022 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1447 0.1220 0.1301 | · | · | · |
| W7 | 1.89× | 0.00225 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00225 0.00225 0.00225 | · | · | · |
| W8 | 2.25× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00225 0.00225 0.00225 | · | · | · |
| W9 | 1.67× | 0.00225 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00225 0.00225 0.00225 | · | · | · |
GeluAndMulFwd¶
x: [M, 28672]
- W1dtype=bf16M=1
ffn-gelu-decode - W2dtype=bf16M=2048
ffn-gelu-prefill - W3dtype=f16M=2048
ffn-gelu-prefill
dtype=f16
- W4x: [17, 514]
tail-fp16
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.71× | 0.0035 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0168 0.0123 | · | · | · |
| W2 | 2.72× | 0.2345 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.6378 0.8785 | · | · | · |
| W3 | 1.38× | 0.2360 | tilelang-ascend tilelang-ascend basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.3236 0.6368 0.8706 | · | · | · |
| W4 | 1.05× | 0.00475 | ops-nn-ew:aclnnGeluMul open-source basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.0035 0.0160 0.0132 | · | · | · |
GeluTanhAndMulFwd¶
x: [M, 28672]
- W1dtype=bf16M=1
ffn-gelu-tanh-decode - W2dtype=bf16M=2048
ffn-gelu-tanh-prefill - W3dtype=f16M=2048
ffn-gelu-tanh-prefill
dtype=f16
- W4x: [17, 514]
tail-fp16
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.48× | 0.00363 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0163 0.0119 | · | · | · |
| W2 | 2.75× | 0.2331 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.6395 0.8731 | · | · | · |
| W3 | 1.37× | 0.2355 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.3174 0.6332 0.8704 | · | · | · |
| W4 | 1.05× | 0.00475 | ops-nn-ew:aclnnGeluMul open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.0035 0.0158 0.0132 | · | · | · |
AcosFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.08× | 0.2213 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.4610 0.4521 | · | · | · |
| W2 | 2.01× | 0.2198 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.4410 0.4290 | · | · | · |
| W3 | 1.90× | 0.2280 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.4323 0.4256 | · | · | · |
| W4 | 2.11× | 3.4029 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 7.1878 7.1807 | · | · | · |
| W5 | 2.03× | 3.3625 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 6.8243 6.8137 | · | · | · |
MaskedFillFwd¶
- W1input: [4096, 4096]mask: [4096, 4096], booldtype=bf16
elementwise-16M - W2input: [4096, 4096]mask: [4096, 4096], booldtype=f16
elementwise-16M - W3input: [4096, 4096]mask: [4096, 4096], booldtype=f32
elementwise-16M - W4input: [16384, 16384]mask: [16384, 16384], booldtype=bf16
elementwise-256M - W5input: [16384, 16384]mask: [16384, 16384], booldtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.74× | 0.0798 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1385 0.1249 | · | · | · |
| W2 | 1.38× | 0.0819 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1091 0.0945 | · | · | · |
| W3 | 1.45× | 0.1338 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1953 0.1800 | · | · | · |
| W4 | 2.17× | 1.1293 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 2.4486 2.4352 | · | · | · |
| W5 | 1.85× | 1.1447 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 2.1277 2.1126 | · | · | · |
AsinFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.76× | 0.2170 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.3822 0.3665 | · | · | · |
| W2 | 1.47× | 0.2110 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.3098 0.3028 | · | · | · |
| W3 | 1.31× | 0.2218 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2898 0.2831 | · | · | · |
| W4 | 2.18× | 3.2824 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 7.1515 7.1376 | · | · | · |
| W5 | 1.48× | 3.2365 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 4.7953 4.7877 | · | · | · |
CompareFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.86× | 0.0583 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0960 0.0963 0.0960 | · | · | · |
| W2 | 2.34× | 0.0470 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1103 0.1419 0.0983 | · | · | · |
| W3 | 1.06× | 0.0794 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0776 0.0780 0.0775 | · | · | · |
| W4 | 1.81× | 0.0530 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0973 0.0973 0.0829 | · | · | · |
| W5 | 1.91× | 0.0500 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0965 0.0935 0.0815 | · | · | · |
| W6 | 1.04× | 0.0922 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0953 0.0925 0.0900 | · | · | · |
HardsigmoidFwd¶
- W1input: [32, 240, 1, 1]dtype=bf16
mbv3-se-gate - W2input: [32, 960, 1, 1]dtype=bf16
mbv3-se-gate-deep - W3input: [32, 960, 1, 1]dtype=f16
mbv3-se-gate-deep - W4input: [32, 240, 1, 1]dtype=f16
mbv3-se-gate
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.12× | 0.002 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.00225 0.0015 | · | · | · |
| W2 | 2.09× | 0.00275 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.00575 0.00225 | · | · | · |
| W3 | 2.00× | 0.00275 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 | 0.00575 0.00225 0.0163 | · | · | · |
| W4 | 1.12× | 0.002 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 | 0.00225 0.0015 0.00487 | · | · | · |
CountNonzeroFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
sparsity-hidden - W3x: [2048, 4096]dtype=f16
sparsity-hidden - W4x: [32, 32768]dtype=f16
sparsity-seq
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.03× | 0.0200 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0803 0.0590 | · | · | · |
| W2 | 1.12× | 0.1025 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1145 0.0885 | · | · | · |
| W3 | 0.84× | 0.1014 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0848 0.0622 | · | · | · |
| W4 | 1.27× | 0.0314 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0395 0.0234 | · | · | · |
LogicalAndFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bool
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W4input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W5input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=bool
hidden-state-prefill - W7input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W8input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.85× | 0.0602 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1025 0.1026 0.1025 | · | · | · |
| W2 | 1.07× | 0.0435 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0357 0.0357 0.0357 | · | · | · |
| W3 | 1.87× | 0.0525 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0820 0.0820 0.0819 | · | · | · |
| W4 | 1.47× | 0.0938 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1324 0.1328 0.1326 | · | · | · |
| W5 | 1.71× | 0.0583 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0990 0.0983 0.0841 | · | · | · |
| W6 | 0.92× | 0.0376 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0348 0.0318 0.0300 | · | · | · |
| W7 | 1.46× | 0.0576 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0830 0.0793 0.0683 | · | · | · |
| W8 | 1.36× | 0.1025 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1384 0.1373 0.1227 | · | · | · |
WhereFwd¶
- W1condition: [16, 256, 56, 56], boolinput: [256, 1, 1]other: [16, 256, 56, 56]dtype=bf16
broadcast - W2condition: [16, 256, 56, 56], boolinput: [256, 1, 1]other: [16, 256, 56, 56]dtype=f16
broadcast - W3condition: [16, 256, 56, 56], boolinput: [256, 1, 1]other: [16, 256, 56, 56]dtype=f32
broadcast - W15condition: [2, 3, 1], boolinput: [1, 3, 5]other: [2, 1, 5]dtype=bf16
three-way-broadcast - W16condition: [2, 3, 1], boolinput: [1, 3, 5]other: [2, 1, 5]dtype=f16
three-way-broadcast - W17condition: [2, 3, 1], boolinput: [1, 3, 5]other: [2, 1, 5]dtype=f32
three-way-broadcast
- W4input: [4096, 4096]dtype=bf16
elementwise-16M - W5input: [4096, 4096]dtype=f16
elementwise-16M - W6input: [4096, 4096]dtype=f32
elementwise-16M - W7input: [16384, 16384]dtype=bf16
elementwise-256M - W8input: [16384, 16384]dtype=f16
elementwise-256M
- W9condition: [4097], boolinput, other: [4097]dtype=bf16
nondiv-tail - W10condition: [4097], boolinput, other: [4097]dtype=f16
nondiv-tail - W11condition: [4097], boolinput, other: [4097]dtype=f32
nondiv-tail - W12condition: [2048, 4096], boolinput, other: [2048, 4096]dtype=bf16
t111-probe-8M - W13condition: [2048, 4096], boolinput, other: [2048, 4096]dtype=f16
t111-probe-8M - W14condition: [2048, 4096], boolinput, other: [2048, 4096]dtype=f32
t111-probe-8M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.79× | 0.0674 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1113 0.1115 0.1110 | · | · | · |
| W2 | 1.14× | 0.0668 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0701 0.0703 0.0700 | · | · | · |
| W3 | 1.21× | 0.1120 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1253 0.1253 0.1253 | · | · | · |
| W4 | 1.28× | 0.1242 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1530 0.1530 0.1527 | · | · | · |
| W5 | 1.12× | 0.1227 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1358 0.1242 0.1320 | · | · | · |
| W6 | 1.10× | 0.2169 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.2404 0.2405 0.2405 | · | · | · |
| W7 | 1.25× | 1.6289 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 2.0532 2.0526 2.0539 | · | · | · |
| W8 | 1.10× | 1.6356 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 1.7980 1.7977 1.7984 | · | · | · |
| W9 | 1.25× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.002 0.002 0.00187 | · | · | · |
| W10 | 1.00× | 0.00225 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00175 0.00175 0.00175 | · | · | · |
| W11 | 1.22× | 0.00225 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.002 0.002 0.002 | · | · | · |
| W12 | 1.28× | 0.0663 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0835 0.0755 0.0793 | · | · | · |
| W13 | 1.11× | 0.0654 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0707 0.0660 0.0675 | · | · | · |
| W14 | 1.09× | 0.1153 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1253 0.1148 0.1207 | · | · | · |
| W15 | 3.50× | 0.0035 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0120 0.0121 0.0120 | · | · | · |
| W16 | 2.86× | 0.0035 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00975 0.0095 0.0095 | · | · | · |
| W17 | 2.35× | 0.00425 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.00988 0.0100 0.0100 | · | · | · |
AnyFwd¶
dtype=bool
- W1x: [4, 128, 4096]dim=[0,2]
3d-multidim-reduce - W2x: [32, 32768]
mask-validation-32k - W3x: [32, 4096]
mask-validation-4k
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.58× | 0.0180 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0310 0.0135 | · | · | · |
| W2 | 1.17× | 0.0152 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0173 0.00725 | · | · | · |
| W3 | 1.43× | 0.0105 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0147 0.00675 | · | · | · |
MishFwd¶
- W1input: [16, 256, 80, 80]dtype=bf16
yolo-p3 - W2input: [16, 256, 80, 80]dtype=f16
yolo-p3 - W3input: [16, 512, 40, 40]dtype=bf16
yolo-p4 - W4input: [16, 512, 40, 40]dtype=f16
yolo-p4
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.37× | 0.1966 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:stage4_opt tilelang-mlir-ascend ⓘ eager · n=5 | 0.2682 0.2607 0.0804 | · | · | · |
| W2 | 1.37× | 0.1955 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:stage4_opt tilelang-mlir-ascend ⓘ eager · n=5 | 0.2667 0.2597 0.0892 | · | · | · |
| W3 | 1.35× | 0.1041 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:stage4_opt tilelang-mlir-ascend ⓘ eager · n=5 | 0.1401 0.1321 0.0470 | · | · | · |
| W4 | 1.33× | 0.1047 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:stage4_opt tilelang-mlir-ascend ⓘ eager · n=5 | 0.1393 0.1315 0.0465 | · | · | · |
LerpTensorFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.37× | 0.1598 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2165 0.2087 | · | · | · |
| W2 | 1.35× | 0.1615 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2160 0.2086 | · | · | · |
| W3 | 1.15× | 0.2800 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.3226 0.3190 | · | · | · |
| W4 | 1.43× | 2.0451 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 2.9049 2.8920 | · | · | · |
| W5 | 1.37× | 2.1231 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 2.9062 2.9065 | · | · | · |
AllFwd¶
dtype=bool
- W1x: [4, 128, 4096]dim=[0,2]
3d-multidim-reduce - W2x: [32, 32768]
mask-validation-32k - W3x: [32, 4096]
mask-validation-4k
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.61× | 0.0175 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0302 0.0130 | · | · | · |
| W2 | 1.02× | 0.0158 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0155 0.00725 | · | · | · |
| W3 | 1.29× | 0.0105 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0140 0.0065 | · | · | · |
PowFwd¶
- W1input: [16, 256, 56, 56]exponent: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]exponent: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]exponent: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, exponent: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, exponent: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, exponent: [2048, 4096]dtype=f32
hidden-state-prefill
- W7exponent, input: [4097]dtype=bf16
nondiv-tail - W8exponent, input: [4097]dtype=f16
nondiv-tail - W9exponent, input: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.76× | 0.2057 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.3535 0.3540 0.3535 | · | · | · |
| W2 | 1.65× | 0.2188 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.3605 0.2831 0.3518 | · | · | · |
| W3 | 1.78× | 0.2040 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.3463 0.3461 0.3463 | · | · | · |
| W4 | 1.49× | 0.1610 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.2402 0.2325 0.2313 | · | · | · |
| W5 | 1.09× | 0.1810 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1857 0.2436 0.2334 0.2310 | · | · | · |
| W6 | 1.05× | 0.1358 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1384 0.2489 0.2385 0.2380 | · | · | · |
| W7 | 1.11× | 0.0045 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.003 0.003 0.003 | · | · | · |
| W8 | 0.78× | 0.00675 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.003 0.003 0.003 | · | · | · |
| W9 | 1.11× | 0.0045 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.003 0.003 0.00287 | · | · | · |
CosFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.30× | 0.2013 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.2607 0.2542 | · | · | · |
| W2 | 1.24× | 0.2077 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2640 0.2585 0.2525 | · | · | · |
| W3 | 1.03× | 0.2037 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.2025 0.2274 0.2223 | · | · | · |
| W4 | 1.34× | 2.9975 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 4.0245 4.0183 | · | · | · |
| W5 | 1.29× | 3.1085 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 4.1885 3.9965 3.9905 | · | · | · |
RemainderFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill - W7input, other: [4097]dtype=bf16
nondiv-tail - W8input, other: [4097]dtype=f16
nondiv-tail - W9input, other: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.32× | 0.1630 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2150 0.2023 | · | · | · |
| W2 | 1.06× | 0.1620 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1725 0.1588 | · | · | · |
| W3 | 1.24× | 0.1640 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2037 0.1886 | · | · | · |
| W4 | 1.28× | 0.1100 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1410 0.1263 | · | · | · |
| W5 | 1.06× | 0.1108 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1172 0.1045 | · | · | · |
| W6 | 1.14× | 0.1278 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1452 0.1340 | · | · | · |
| W7 | 1.20× | 0.00313 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.004 0.00225 | · | · | · |
| W8 | 1.04× | 0.00325 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.00325 0.00213 | · | · | · |
| W9 | 1.15× | 0.00325 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.002 0.002 | · | · | · |
ClampFwd¶
- W1input, min, max: [4096, 4096]dtype=bf16
elementwise-16M - W2input, min, max: [4096, 4096]dtype=f16
elementwise-16M - W3input, min, max: [4096, 4096]dtype=f32
elementwise-16M - W10input, min, max: [16384, 16384]dtype=bf16
elementwise-256M - W11input, min, max: [16384, 16384]dtype=f16
elementwise-256M
- W4input, max: [4096, 4096]dtype=bf16
elementwise-16M-max-only - W5input, max: [4096, 4096]dtype=f16
elementwise-16M-max-only - W6input, max: [4096, 4096]dtype=f32
elementwise-16M-max-only - W12input, max: [16384, 16384]dtype=bf16
elementwise-256M-max-only - W13input, max: [16384, 16384]dtype=f16
elementwise-256M-max-only
- W7input, min: [4096, 4096]dtype=bf16
elementwise-16M-min-only - W8input, min: [4096, 4096]dtype=f16
elementwise-16M-min-only - W9input, min: [4096, 4096]dtype=f32
elementwise-16M-min-only - W14input, min: [16384, 16384]dtype=bf16
elementwise-256M-min-only - W15input, min: [16384, 16384]dtype=f16
elementwise-256M-min-only
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.09× | 0.1575 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1703 0.1656 | · | · | · |
| W2 | 0.95× | 0.1591 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1504 0.1491 | · | · | · |
| W3 | 1.01× | 0.2770 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.2807 0.2815 | · | · | · |
| W4 | 1.39× | 0.1059 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1470 0.1422 | · | · | · |
| W5 | 1.11× | 0.1060 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1071 0.1070 | · | · | · |
| W6 | 1.06× | 0.2009 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.2100 0.2102 | · | · | · |
| W7 | 1.53× | 0.1062 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1489 0.1462 | · | · | · |
| W8 | 1.10× | 0.1085 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1185 0.1077 | · | · | · |
| W9 | 1.07× | 0.2045 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.2190 0.2115 | · | · | · |
| W10 | 1.05× | 2.0339 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 2.1299 2.1202 | · | · | · |
| W11 | 0.96× | 2.0513 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 1.9509 1.9507 | · | · | · |
| W12 | 1.47× | 1.4291 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 1.8487 1.9780 | · | · | · |
| W13 | 1.07× | 1.4126 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 1.5100 1.5101 | · | · | · |
| W14 | 1.50× | 1.3942 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 2.1264 2.2654 | · | · | · |
| W15 | 1.07× | 1.4265 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 1.5146 1.5129 | · | · | · |
TanhFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.28× | 0.0818 | ops-nn-ew-r273:aclnnForeachTanh open-source basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.1030 0.1573 0.1495 | · | · | · |
| W2 | 1.01× | 0.0665 | ops-nn-ew-r273:aclnnForeachTanh open-source basistilelang-ascend tilelang-ascendtorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.0645 0.2742 0.1557 0.1490 | · | · | · |
| W3 | 1.06× | 0.1182 | ops-nn-ew-r273:aclnnForeachTanh open-source basistilelang-ascend tilelang-ascendtorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1172 0.2421 0.1475 0.1409 | · | · | · |
| W4 | 1.28× | 1.1625 | ops-nn-ew-r273:aclnnForeachTanh open-source basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 1.4836 2.3478 2.3407 | · | · | · |
| W5 | 0.96× | 0.9285 | ops-nn-ew-r273:aclnnForeachTanh open-source basistilelang-ascend tilelang-ascendtorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.8892 3.7560 2.3360 2.3301 | · | · | · |
DivFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill - W7input, other: [4097]dtype=bf16
nondiv-tail - W8input, other: [4097]dtype=f16
nondiv-tail - W9input, other: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.00× | 0.0679 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.0630 0.0630 0.0633 | · | · | · |
| W2 | 1.49× | 0.0485 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0726 0.1713 0.0658 | · | · | · |
| W3 | 1.11× | 0.0877 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0912 0.0912 0.0912 | · | · | · |
| W4 | 1.10× | 0.0578 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0622 0.0587 0.0583 | · | · | · |
| W5 | 1.07× | 0.0568 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0765 0.0611 0.0585 0.0583 | · | · | · |
| W6 | 1.02× | 0.1020 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0990 0.1037 0.1040 0.1016 | · | · | · |
| W7 | 1.00× | 0.00225 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
| W8 | 1.12× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
| W9 | 1.00× | 0.00225 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
SinFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.46× | 0.1698 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.2487 0.2422 | · | · | · |
| W2 | 0.90× | 0.1777 | ops-nn-ew-r273:aclnnForeachSin open-source basistilelang-ascend tilelang-ascendtorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1593 0.2239 0.2467 0.2407 | · | · | · |
| W3 | 0.86× | 0.1762 | ops-nn-ew-r273:aclnnForeachSin open-source basistilelang-ascend tilelang-ascendtorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1457 0.1755 0.2157 0.2102 | · | · | · |
| W4 | 1.51× | 2.5425 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 3.8312 3.8260 | · | · | · |
| W5 | 0.91× | 2.6194 | ops-nn-ew-r273:aclnnForeachSin open-source basistilelang-ascend tilelang-ascendtorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 2.3805 3.5980 3.8037 3.7988 | · | · | · |
SiluAndMulFwd¶
x: [M, 28672]
- W1dtype=bf16M=1
llama-8b-swiglu-decode - W2dtype=bf16M=2048
llama-8b-swiglu-prefill - W3dtype=f16M=2048
llama-8b-swiglu-prefill
dtype=f16
- W4x: [17, 514]
tail-fp16
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.58× | 0.003 | ops-nn-ew:aclnnSwiGlu open-source basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.00475 0.0168 0.0115 | · | · | · |
| W2 | 0.93× | 0.1994 | ops-nn-ew:aclnnSwiGlu open-source basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.1802 0.6261 0.8649 | · | · | · |
| W3 | 0.91× | 0.2039 | ops-nn-ew:aclnnSwiGlu open-source basistilelang-ascend tilelang-ascendtorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.1815 0.3065 0.6272 0.8635 | · | · | · |
| W4 | 1.00× | 0.00425 | ops-nn-ew:aclnnSwiGlu open-source basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.00325 0.0152 0.0130 | · | · | · |
AddFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.12× | 0.0609 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0679 0.0626 | · | · | · |
| W2 | 1.18× | 0.0506 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0587 0.0485 | · | · | · |
| W3 | 1.10× | 0.0885 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0983 0.0916 | · | · | · |
| W4 | 1.10× | 0.0563 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:example tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0620 0.0578 0.2890 | · | · | · |
| W5 | 0.96× | 0.0553 | tilelang-ascend tilelang-ascend basistorch_compile_aclgraph torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0500 0.0540 0.0504 0.1536 | · | · | · |
| W6 | 0.99× | 0.1013 | tilelang-ascend tilelang-ascend basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.0971 0.1035 0.1010 | · | · | · |
SubFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill - W7input, other: [4097]dtype=bf16
nondiv-tail - W8input, other: [4097]dtype=f16
nondiv-tail - W9input, other: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.09× | 0.0624 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.0615 0.0617 0.0616 | · | · | · |
| W2 | 1.11× | 0.0480 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0530 0.0610 0.0471 | · | · | · |
| W3 | 1.11× | 0.0872 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0897 0.0897 0.0897 | · | · | · |
| W4 | 1.09× | 0.0545 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0590 0.0563 0.0568 | · | · | · |
| W5 | 1.04× | 0.0521 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0737 0.0546 0.0542 0.0505 | · | · | · |
| W6 | 1.03× | 0.1008 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0985 0.1051 0.1030 0.1010 | · | · | · |
| W7 | 1.12× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
| W8 | 0.88× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
| W9 | 1.00× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
LogicalNotFwd¶
- W1input: [4096, 4096]dtype=bool
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bool
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.80× | 0.0435 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0352 0.0305 | · | · | · |
| W2 | 1.32× | 0.0615 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0818 0.0694 | · | · | · |
| W3 | 1.35× | 0.0998 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1355 0.1196 | · | · | · |
| W4 | 0.81× | 0.5570 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.4525 0.4495 | · | · | · |
MulFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill - W7input, other: [4097]dtype=bf16
nondiv-tail - W8input, other: [4097]dtype=f16
nondiv-tail - W9input, other: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.13× | 0.0607 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0622 0.0617 0.0620 | · | · | · |
| W2 | 1.10× | 0.0498 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0536 0.0517 0.0483 | · | · | · |
| W3 | 1.08× | 0.0887 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0892 0.0887 0.0890 | · | · | · |
| W4 | 1.10× | 0.0560 | ops-nn-ew:aclnnForeachMulList open-source ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0583 0.0607 0.0570 0.0536 | · | · | · |
| W5 | 1.01× | 0.0540 | ops-nn-ew:aclnnForeachMulList open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0561 0.0761 0.0548 0.0550 0.0515 | · | · | · |
| W6 | 1.03× | 0.1000 | ops-nn-ew:aclnnForeachMulList open-source ⓘtilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1032 0.0968 0.1040 0.1017 0.1010 | · | · | · |
| W7 | 0.89× | 0.00225 | ops-nn-ew:aclnnForeachMulList open-source ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.009 0.0015 0.0015 0.0015 | · | · | · |
| W8 | 0.88× | 0.002 | ops-nn-ew:aclnnForeachMulList open-source ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.00887 0.0015 0.0015 0.0015 | · | · | · |
| W9 | 1.00× | 0.002 | ops-nn-ew:aclnnForeachMulList open-source ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.00875 0.0015 0.0015 0.0015 | · | · | · |
Log1pFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.06× | 0.0675 | ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0670 0.1060 0.1067 0.0990 | · | · | · |
| W2 | 1.07× | 0.0620 | ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0626 0.1056 0.1062 0.0988 | · | · | · |
| W3 | 0.99× | 0.1151 | ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1151 0.1378 0.1355 0.1310 | · | · | · |
| W4 | 0.94× | 0.9560 | ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.8952 1.5397 1.5394 1.5395 | · | · | · |
| W5 | 1.02× | 0.8628 | ops-nn-ew-r273:aclnnForeachLog1p open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.8782 1.5441 1.5148 1.5386 | · | · | · |
ReciprocalFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.91× | 0.0685 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0625 0.0584 | · | · | · |
| W2 | 0.97× | 0.0665 | tilelang-ascend tilelang-ascend ⓘtorch_npu eager vendor library basis | 0.0915 0.0590 | · | · | · |
| W3 | 0.98× | 0.1172 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1128 0.1154 0.1115 | · | · | · |
| W4 | 0.92× | 0.9778 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.9024 0.9036 | · | · | · |
| W5 | 1.03× | 0.8831 | tilelang-ascend tilelang-ascend ⓘtorch_npu eager vendor library basis | 1.4269 0.8991 | · | · | · |
SignFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.96× | 0.0755 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0725 0.0643 | · | · | · |
| W2 | 0.96× | 0.0700 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0664 0.0578 | · | · | · |
| W3 | 0.92× | 0.1260 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1155 0.1092 | · | · | · |
| W4 | 0.95× | 1.0490 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.9931 0.9901 | · | · | · |
| W5 | 0.93× | 0.9585 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8985 0.8812 | · | · | · |
Expm1Fwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.86× | 0.0680 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0580 0.0553 | · | · | · |
| W2 | 0.92× | 0.0630 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0580 0.0560 | · | · | · |
| W3 | 0.98× | 0.1143 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1103 0.1077 | · | · | · |
| W4 | 0.92× | 0.9639 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.8835 0.8862 | · | · | · |
| W5 | 1.04× | 0.8598 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8889 0.8852 | · | · | · |
SqrtFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.88× | 0.0707 | ops-nn-ew:aclnnForeachSqrt open-sourcetorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0669 0.0633 0.0583 | · | · | · |
| W2 | 0.96× | 0.0599 | ops-nn-ew:aclnnForeachSqrt open-sourcetilelang-ascend tilelang-ascendtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0622 0.0558 0.0610 0.0535 | · | · | · |
| W3 | 0.98× | 0.1144 | ops-nn-ew:aclnnForeachSqrt open-sourcetilelang-ascend tilelang-ascend basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1155 0.1085 0.1116 0.1090 | · | · | · |
| W4 | 0.93× | 0.9390 | ops-nn-ew:aclnnForeachSqrt open-sourcetorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.8890 0.8855 0.8849 | · | · | · |
| W5 | 0.95× | 0.8799 | ops-nn-ew:aclnnForeachSqrt open-sourcetilelang-ascend tilelang-ascend basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.8722 0.8336 0.8814 0.8766 | · | · | · |
LogFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.88× | 0.0700 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0668 0.0576 | · | · | · |
| W2 | 0.98× | 0.0639 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.0583 0.0615 0.0587 | · | · | · |
| W3 | 0.99× | 0.1153 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1106 0.1128 0.1119 | · | · | · |
| W4 | 0.92× | 0.9515 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.8780 0.8821 | · | · | · |
| W5 | 0.93× | 0.9524 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.8526 0.8729 0.8776 | · | · | · |
AbsFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.88× | 0.0668 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0573 0.0540 | · | · | · |
| W2 | 0.98× | 0.0602 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0530 0.0553 0.0535 0.1972 | · | · | · |
| W3 | 0.97× | 0.1135 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.1060 0.1087 0.1065 | · | · | · |
| W4 | 0.93× | 0.9407 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.8788 0.8785 | · | · | · |
| W5 | 0.94× | 0.9439 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.8436 0.8655 0.8608 3.4259 | · | · | · |
NegFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.85× | 0.0705 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0609 0.0565 | · | · | · |
| W2 | 0.93× | 0.0610 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0563 0.0535 | · | · | · |
| W3 | 0.98× | 0.1150 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1118 0.1098 | · | · | · |
| W4 | 0.91× | 0.9720 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.8804 0.8830 | · | · | · |
| W5 | 1.01× | 0.8565 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8742 0.8632 | · | · | · |
ExpFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.88× | 0.0683 | ops-nn-ew:aclnnForeachExp open-source ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0650 0.0620 0.0597 0.0548 | · | · | · |
| W2 | 0.94× | 0.0639 | ops-nn-ew:aclnnForeachExp open-source ⓘtilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0631 0.0550 0.0600 0.0624 0.0578 | · | · | · |
| W3 | 0.97× | 0.1141 | ops-nn-ew:aclnnForeachExp open-source ⓘtilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1145 0.1065 0.1094 0.1105 0.1070 | · | · | · |
| W4 | 0.93× | 0.9497 | ops-nn-ew:aclnnForeachExp open-source ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.8879 0.8811 0.8811 0.8811 | · | · | · |
| W5 | 0.94× | 0.8901 | ops-nn-ew:aclnnForeachExp open-source ⓘtilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.8778 0.8307 0.8845 0.8781 0.8760 | · | · | · |
Log2Fwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.86× | 0.0665 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0587 0.0546 | · | · | · |
| W2 | 0.94× | 0.0607 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0590 0.0534 | · | · | · |
| W3 | 0.97× | 0.1172 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1135 0.1098 | · | · | · |
| W4 | 0.90× | 0.9738 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.8884 0.8770 | · | · | · |
| W5 | 0.98× | 0.8815 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.8756 0.8696 | · | · | · |
SigmoidFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.82× | 0.0760 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0620 0.0574 | · | · | · |
| W2 | 0.98× | 0.0628 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0885 0.0611 0.0555 0.2163 | · | · | · |
| W3 | 0.97× | 0.1150 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1119 0.1111 0.1082 | · | · | · |
| W4 | 0.85× | 1.0755 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.9140 0.9096 | · | · | · |
| W5 | 1.01× | 0.9009 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 1.4780 0.9135 0.9025 3.7019 | · | · | · |
ReluFwd¶
- W1input: [1, 4096]dtype=bf16
hidden-state-decode - W2input: [2048, 4096]dtype=bf16
hidden-state-prefill - W3input: [2048, 4096]dtype=f16
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.00× | 0.00175 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.00175 0.00175 0.00125 0.00325 | · | · | · |
| W2 | 0.88× | 0.0395 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0345 0.0338 0.0297 0.0285 | · | · | · |
| W3 | 0.86× | 0.0400 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0290 0.0323 0.0350 0.0290 0.0287 | · | · | · |
SiluFwd¶
- W1input: [1, 14336]dtype=bf16
llama-8b-ffn-decode - W2input: [2048, 14336]dtype=bf16
llama-8b-ffn-prefill - W3input: [2048, 14336]dtype=f16
llama-8b-ffn-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.18× | 0.00275 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.00325 0.0035 0.00175 | · | · | · |
| W2 | 0.78× | 0.1406 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.1077 0.1071 0.1030 | · | · | · |
| W3 | 0.77× | 0.1385 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.2223 0.1069 0.1062 0.1016 0.3487 | · | · | · |
RsqrtFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.89× | 0.0710 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0636 0.0583 | · | · | · |
| W2 | 0.86× | 0.0705 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.0542 0.0612 0.0574 | · | · | · |
| W3 | 0.91× | 0.1225 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1080 0.1128 0.1091 | · | · | · |
| W4 | 0.90× | 1.0069 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.9074 0.9107 | · | · | · |
| W5 | 0.88× | 1.0249 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.8526 0.9008 0.9019 | · | · | · |
RoundFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.85× | 0.0715 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0607 0.0565 | · | · | · |
| W2 | 0.84× | 0.0700 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0628 0.0553 | · | · | · |
| W3 | 0.91× | 0.1236 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1130 0.1085 | · | · | · |
| W4 | 0.88× | 1.0065 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.8838 0.8862 | · | · | · |
| W5 | 0.90× | 0.9815 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.8799 0.8858 | · | · | · |
TruncFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.83× | 0.0741 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0622 0.0573 | · | · | · |
| W2 | 0.86× | 0.0704 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0602 0.0558 | · | · | · |
| W3 | 0.90× | 0.1242 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1118 0.1098 | · | · | · |
| W4 | 0.88× | 1.0052 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.8825 0.8858 | · | · | · |
| W5 | 0.89× | 0.9839 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.8751 0.8851 | · | · | · |
FloorFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.82× | 0.0722 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0592 0.0550 | · | · | · |
| W2 | 0.83× | 0.0698 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0578 0.0542 | · | · | · |
| W3 | 0.91× | 0.1226 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1115 0.1085 | · | · | · |
| W4 | 0.88× | 1.0051 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8828 0.8791 | · | · | · |
| W5 | 0.89× | 0.9900 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8868 0.8824 | · | · | · |
CeilFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.82× | 0.0719 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0592 0.0545 | · | · | · |
| W2 | 0.83× | 0.0693 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0580 0.0537 | · | · | · |
| W3 | 0.90× | 0.1232 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1113 0.1087 | · | · | · |
| W4 | 0.88× | 1.0037 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8821 0.8775 | · | · | · |
| W5 | 0.89× | 0.9910 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.8888 0.8819 | · | · | · |
HardswishFwd¶
- W1input: [32, 96, 56, 56]dtype=bf16
mbv3-stage2 - W2input: [32, 96, 56, 56]dtype=f16
mbv3-stage2 - W3input: [32, 240, 28, 28]dtype=bf16
mbv3-stage3 - W4input: [32, 240, 28, 28]dtype=f16
mbv3-stage3
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.89× | 0.0445 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0395 0.0350 | · | · | · |
| W2 | 0.84× | 0.0465 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0398 0.0352 | · | · | · |
| W3 | 0.81× | 0.0330 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0272 0.0230 | · | · | · |
| W4 | 0.82× | 0.0333 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0275 0.0230 | · | · | · |
BmmFwd¶
a: [B, M, K]b: [B, K, N]
- W1dtype=bf16B=64M=128K=2048N=128
mha-decode-b64-pv - W2dtype=f16B=64M=128K=2048N=128
mha-decode-b64-pv - W3dtype=bf16B=64M=128K=128N=2048
mha-decode-b64-qk - W4dtype=f16B=64M=128K=128N=2048
mha-decode-b64-qk - W5dtype=bf16B=128M=512K=2048N=512
moe-prefill-b128 - W6dtype=bf16B=8M=128K=128N=128
small-b8-128 - W7dtype=f16B=8M=128K=128N=128
small-b8-128 - W8dtype=bf16B=16M=512K=512N=512
square-b16-512 - W9dtype=f16B=16M=512K=512N=512
square-b16-512 - W10dtype=bf16B=32M=256K=256N=256
square-b32-256 - W11dtype=f16B=32M=256K=256N=256
square-b32-256 - W12dtype=bf16B=4M=4096K=4096N=4096
square-b4-4k - W13dtype=bf16B=8M=1024K=1024N=1024
square-b8-1k - W14dtype=f16B=8M=1024K=1024N=1024
square-b8-1k - W15dtype=bf16B=8M=2048K=2048N=2048
square-b8-2k - W16dtype=f16B=8M=2048K=2048N=2048
square-b8-2k
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.78× | 0.1376 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1010 0.1049 0.1295 0.1045 0.1046 0.1046 | · | · | · |
| W2 | 0.76× | 0.1400 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1014 0.1055 0.1313 0.1052 0.1055 0.1052 | · | · | · |
| W3 | 0.78× | 0.0926 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0653 0.0675 0.0899 0.0691 0.0690 0.0691 | · | · | · |
| W4 | 0.76× | 0.0930 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0650 0.0660 0.0897 0.0668 0.0668 0.0668 | · | · | · |
| W5 | 0.99× | 0.7890 | catlass-r252:r252_bmm open-source ⓘops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.8089 0.7823 1.5370 0.7802 0.7795 0.7782 | · | · | · |
| W6 | 0.56× | 0.0103 | catlass-r252:r252_bmm open-source ⓘops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.00675 0.00825 0.00375 0.0120 0.0055 0.0085 | · | · | · |
| W7 | 0.59× | 0.0103 | catlass-r252:r252_bmm open-source ⓘops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0065 0.00925 0.00375 0.0150 0.0055 0.00925 | · | · | · |
| W8 | 0.96× | 0.0500 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0395 0.0488 0.0602 0.0475 0.0475 0.0475 | · | · | · |
| W9 | 0.90× | 0.0535 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0420 0.0490 0.0614 0.0490 0.0489 0.0490 | · | · | · |
| W10 | 0.82× | 0.0307 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0185 0.0248 0.0227 0.0255 0.0255 0.0255 | · | · | · |
| W11 | 0.85× | 0.0293 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0187 0.0243 0.0227 0.0245 0.0246 0.0245 | · | · | · |
| W12 | 0.96× | 1.8649 | catlass-r252:r252_bmm open-source ⓘops-nn-gemm:aclnnBatchMatMul open-source ⓘ basistilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 1.9143 1.7851 5.2805 1.7854 1.7873 1.7857 | · | · | · |
| W13 | 0.89× | 0.1126 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0917 0.1101 0.1956 0.1105 0.0925 0.1103 | · | · | · |
| W14 | 0.87× | 0.1110 | catlass-r252:r252_bmm open-source ⓘ basisops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0902 0.1096 0.1953 0.1120 0.0917 0.1095 | · | · | · |
| W15 | 0.92× | 0.5693 | catlass-r252:r252_bmm open-source ⓘops-nn-gemm:aclnnBatchMatMul open-source ⓘ basistilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.5575 0.5215 1.3916 0.5226 0.5230 0.5226 | · | · | · |
| W16 | 0.92× | 0.5667 | catlass-r252:r252_bmm open-source ⓘops-nn-gemm:aclnnBatchMatMul open-source ⓘtilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.5573 0.5184 1.3915 0.5179 0.5185 0.5185 | · | · | · |
EluFwd¶
- W1input: [2048, 4096]dtype=bf16
mlp-hidden - W2input: [2048, 4096]dtype=f16
mlp-hidden - W3input: [2048, 8192]dtype=bf16
mlp-hidden-wide - W4input: [2048, 8192]dtype=f16
mlp-hidden-wide
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.78× | 0.0437 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0340 0.0352 0.0312 0.0332 | · | · | · |
| W2 | 0.83× | 0.0419 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0348 0.0352 0.0312 0.0291 | · | · | · |
| W3 | 0.82× | 0.0760 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0625 0.0630 0.0583 | · | · | · |
| W4 | 0.85× | 0.0752 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0630 0.0649 0.0591 0.0581 | · | · | · |
LeakyReluFwd¶
- W1input: [16, 256, 64, 64]dtype=bf16
gan-feat - W2input: [16, 512, 32, 32]dtype=bf16
gan-feat-deep - W3input: [16, 512, 32, 32]dtype=f16
gan-feat-deep - W4input: [16, 256, 64, 64]dtype=f16
gan-feat
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.81× | 0.0710 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0625 0.0550 | · | · | · |
| W2 | 0.79× | 0.0420 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0330 0.0288 | · | · | · |
| W3 | 0.77× | 0.0415 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0312 0.0285 0.0280 | · | · | · |
| W4 | 0.82× | 0.0715 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0587 0.0550 0.0512 | · | · | · |
ErfFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.81× | 0.1830 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1477 0.1427 | · | · | · |
| W2 | 0.81× | 0.1805 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1460 0.1412 | · | · | · |
| W3 | 0.74× | 0.1921 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1415 0.1360 | · | · | · |
| W4 | 0.80× | 2.7886 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 2.2340 2.2291 | · | · | · |
| W5 | 0.80× | 2.7494 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 2.2070 2.2020 | · | · | · |
GeluFwd¶
- W1input: [1, 14336]dtype=bf16
llama-8b-ffn-decode - W2input: [2048, 14336]dtype=bf16
llama-8b-ffn-prefill - W3input: [2048, 14336]dtype=f16
llama-8b-ffn-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.14× | 0.0035 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00375 0.00325 0.00175 | · | · | · |
| W2 | 0.58× | 0.1930 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1123 0.1133 0.1085 | · | · | · |
| W3 | 0.57× | 0.1910 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1085 0.1106 0.1050 | · | · | · |
TanFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.51× | 0.4537 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2295 0.2255 | · | · | · |
| W2 | 0.52× | 0.4397 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2280 0.2231 | · | · | · |
| W3 | 0.46× | 0.4556 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2115 0.2057 | · | · | · |
| W4 | 0.50× | 7.1116 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 3.5642 3.5600 | · | · | · |
| W5 | 0.52× | 6.8262 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 3.5307 3.5265 | · | · | · |
CastFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.40× | 0.1850 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0735 0.0961 | · | · | · |
| W2 | 0.38× | 0.1847 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0703 0.0940 | · | · | · |
| W3 | 0.57× | 0.2052 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1165 0.1121 | · | · | · |
| W4 | 0.44× | 2.8824 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 1.2768 2.5775 | · | · | · |
| W5 | 0.43× | 2.8820 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 1.2505 2.5269 | · | · | · |
Im2col¶
- W1input: [2, 128, 28, 28]bias: [512]dtype=f16C_out=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1-bias - W3input: [2, 512, 28, 28]bias: [128]dtype=f16C_out=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1-bias - W5input: [1, 512, 7, 7]bias: [2048]dtype=f16C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1-bias - W8input: [1, 256, 112, 112]bias: [512]dtype=f16C_out=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1-bias - W10input: [1, 256, 14, 14]bias: [1024]dtype=f16C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1-bias - W12input: [1, 64, 56, 56]bias: [128]dtype=f16C_out=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1-bias - W16input: [2, 64, 56, 56]bias: [256]dtype=bf16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W17input: [2, 64, 56, 56]bias: [256]dtype=f16C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1-bias - W20input: [2, 64, 56, 56]bias: [64]dtype=bf16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W21input: [2, 64, 56, 56]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3-bias - W24input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2-bias - W26input: [1, 128, 56, 56]bias: [256]dtype=f16C_out=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2-bias - W28input: [1, 3, 112, 112]bias: [64]dtype=f16C_out=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2-bias - W31input: [1, 128, 28, 28]bias: [128]dtype=bf16C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2-bias
input: [N, C_in, H, W]
- W2dtype=f16N=2C_in=128H=28W=28C_out=512kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-expand-1x1 - W4dtype=f16N=2C_in=512H=28W=28C_out=128kH=1kW=1stride=[1,1]padding=[0,0]
bottleneck-reduce-1x1 - W6dtype=f16N=1C_in=512H=7W=7C_out=2048kH=1kW=1stride=[1,1]padding=[0,0]
classifier-1x1 - W7dtype=f16N=1C_in=2048H=32W=32C_out=256kH=3kW=3stride=[1,1]padding=[12,12]dilation=[12,12]
deeplabv3-aspp-3x3-rate12 - W9dtype=f16N=1C_in=256H=112W=112C_out=512kH=3kW=3stride=[1,1]padding=[1,1]
highres-3x3-s1 - W11dtype=f16N=1C_in=256H=14W=14C_out=1024kH=1kW=1stride=[1,1]padding=[0,0]
late-stage-1x1 - W13dtype=f16N=1C_in=64H=56W=56C_out=128kH=5kW=5stride=[1,1]padding=[2,2]
midres-5x5-s1 - W14dtype=f16N=1C_in=32H=56W=56C_out=32kH=3kW=3stride=[1,1]padding=[1,1]groups=32
mobilenetv2-depthwise - W15dtype=bf16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W18dtype=f16N=2C_in=64H=56W=56C_out=256kH=1kW=1stride=[1,1]padding=[0,0]
resnet-1x1 - W19dtype=bf16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W22dtype=f16N=2C_in=64H=56W=56C_out=64kH=3kW=3stride=[1,1]padding=[1,1]
resnet-3x3 - W23dtype=f16N=1C_in=128H=28W=28C_out=256kH=3kW=3stride=[1,1]padding=[1,1]groups=32
resnext-grouped-3x3 - W25dtype=f16N=1C_in=128H=56W=56C_out=256kH=3kW=3stride=[2,2]padding=[1,1]
stage-transition-3x3-s2 - W27dtype=f16N=1C_in=128H=56W=56C_out=256kH=5kW=5stride=[2,2]padding=[2,2]
stage-transition-5x5-s2 - W29dtype=f16N=1C_in=3H=112W=112C_out=64kH=3kW=3stride=[2,2]padding=[1,1]
stem-3x3-s2 - W30dtype=bf16N=1C_in=128H=28W=28C_out=128kH=3kW=3stride=[2,2]padding=[1,1]
stride2
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.02× | 0.0100 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00625 0.00625 0.00625 | · | · | · |
| W2 | 1.00× | 0.0103 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00625 0.00625 0.00625 | · | · | · |
| W3 | 0.76× | 0.0235 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0150 0.0110 0.0107 | · | · | · |
| W4 | 0.76× | 0.0235 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0150 0.0107 0.0105 | · | · | · |
| W5 | 0.22× | 0.0730 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0107 0.0109 0.0107 | · | · | · |
| W6 | 0.22× | 0.0732 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0107 0.0107 0.0107 | · | · | · |
| W7 | 0.02× | 7.1800 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1583 0.1537 0.1578 | · | · | · |
| W8 | 0.59× | 0.3965 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.2442 0.2289 0.2300 | · | · | · |
| W9 | 0.59× | 0.3964 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2320 0.2295 | · | · | · |
| W10 | 0.60× | 0.0175 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00637 0.00625 0.00637 | · | · | · |
| W11 | 0.60× | 0.0174 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0065 0.00637 0.00625 | · | · | · |
| W12 | 1.19× | 0.1928 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.2300 0.2281 0.2274 | · | · | · |
| W13 | 1.13× | 0.1948 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.2214 0.2255 | · | · | · |
| W14 | 1.99× | 0.0410 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0739 0.0686 0.0741 | · | · | · |
| W15 | 3.13× | 0.00975 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0219 0.0227 0.0225 | · | · | · |
| W16 | 3.30× | 0.0100 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0249 0.0245 0.0243 | · | · | · |
| W17 | 3.04× | 0.0104 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0240 0.0249 0.0240 | · | · | · |
| W18 | 3.05× | 0.0103 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0236 0.0240 0.0231 | · | · | · |
| W19 | 1.49× | 0.1118 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1594 0.1479 | · | · | · |
| W20 | 1.50× | 0.1110 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1678 0.1370 0.1466 | · | · | · |
| W21 | 1.50× | 0.1103 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1725 0.1704 0.1608 | · | · | · |
| W22 | 1.50× | 0.1100 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1645 0.1555 | · | · | · |
| W23 | 0.76× | 0.0675 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0467 0.0474 0.0460 | · | · | · |
| W24 | 0.02× | 2.8989 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0616 0.0442 0.0449 | · | · | · |
| W25 | 0.02× | 2.8984 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0534 0.0437 | · | · | · |
| W26 | 0.01× | 8.2525 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1316 0.1133 0.1140 | · | · | · |
| W27 | 0.01× | 8.2511 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1242 0.1156 | · | · | · |
| W28 | 0.24× | 0.2868 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0678 0.0670 0.0727 | · | · | · |
| W29 | 0.24× | 0.2864 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.0678 0.0781 | · | · | · |
| W30 | 0.01× | 0.7562 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0105 0.00775 0.0075 | · | · | · |
| W31 | 0.01× | 0.7576 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0105 0.00725 0.00725 | · | · | · |
BiasAddFwd¶
- W1input: [16, 256, 56, 56]bias: [56]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]bias: [56]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]bias: [56]dtype=f32
cnn-feat-broadcast - W4input: [2048, 4096]bias: [4096]dtype=bf16
hidden-state-prefill - W5input: [2048, 4096]bias: [4096]dtype=f16
hidden-state-prefill - W6input: [2048, 4096]bias: [4096]dtype=f32
hidden-state-prefill
- W7bias, input: [4097]dtype=bf16
nondiv-tail - W8bias, input: [4097]dtype=f16
nondiv-tail - W9bias, input: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.01× | 8.1152 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0545 0.0545 0.0542 | · | · | · |
| W2 | 0.01× | 8.1036 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0550 0.0490 0.0481 | · | · | · |
| W3 | 0.09× | 1.0171 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0886 0.0885 0.0887 | · | · | · |
| W4 | 1.03× | 0.0390 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0393 0.0362 0.0350 | · | · | · |
| W5 | 0.96× | 0.0385 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0367 0.0350 0.0305 | · | · | · |
| W6 | 0.98× | 0.0650 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0631 0.0610 0.0585 | · | · | · |
| W7 | 1.00× | 0.00225 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0015 0.0015 0.0015 | · | · | · |
| W8 | 0.88× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0015 0.00125 0.0015 | · | · | · |
| W9 | 1.00× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0015 0.0015 0.0015 | · | · | · |
ProdFwd¶
- W1x: [2048, 4096]dtype=bf16
hidden-state-reduce - W2x: [2048, 4096]dtype=f16
hidden-state-reduce - W3x: [64, 32768]dtype=bf16
long-seq-reduce
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.20× | 0.4052 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0779 0.0650 | · | · | · |
| W2 | 0.20× | 0.4060 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0789 0.0636 | · | · | · |
| W3 | — | · | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0376 0.0194 | · | · | · |
GemvFwd¶
a: [M, K]x: [K]
trans_a=false
- W1dtype=bf16M=128K=2048n=7168trans_b=true
ds-v3-decode-down - W2dtype=bf16M=128K=7168n=2112trans_b=true
ds-v3-decode-gate-up - W3dtype=bf16M=4096K=7168n=4096trans_b=true
ds-v3-prefill-attn-proj - W4dtype=f16M=4096K=7168n=4096trans_b=true
ds-v3-prefill-attn-proj - W5dtype=bf16M=4096K=2048n=7168trans_b=true
ds-v3-prefill-down - W6dtype=bf16M=4096K=7168n=2112trans_b=true
ds-v3-prefill-gate-up - W7dtype=bf16M=4096K=16384n=7168trans_b=true
k-dominant-7168x16384 - W8dtype=bf16M=16K=7168n=4096trans_b=true
mid-m16-attn - W9dtype=bf16M=32K=7168n=4096trans_b=true
mid-m32-attn - W10dtype=bf16M=64K=2048n=7168trans_b=true
mid-m64-down - W11dtype=bf16M=96K=7168n=2112trans_b=true
mid-m96-gate-up - W12dtype=bf16M=1024K=1024n=1024trans_b=false
square-1k-nn - W13dtype=f16M=1024K=1024n=1024trans_b=false
square-1k-nn - W14dtype=bf16M=4096K=1536n=24576trans_b=true
wide-n-24576
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.41× | 0.0777 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0320 0.0213 | · | · | · |
| W2 | 0.14× | 0.2875 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0403 0.0269 | · | · | · |
| W3 | 0.08× | 2.0386 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2107 0.1679 | · | · | · |
| W4 | 0.09× | 1.8646 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1670 0.1635 | · | · | · |
| W5 | 0.09× | 1.0717 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0930 0.0895 | · | · | · |
| W6 | 0.09× | 1.8217 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1708 0.1689 | · | · | · |
| W7 | 0.07× | 3.0372 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.2214 0.2199 0.2201 | · | · | · |
| W8 | 0.10× | 0.2890 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0248 0.0245 0.0249 | · | · | · |
| W9 | 0.14× | 0.2772 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0250 0.0245 0.0245 | · | · | · |
| W10 | 0.32× | 0.0943 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0198 0.0195 0.0196 | · | · | · |
| W11 | 0.11× | 0.3619 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0261 0.0262 0.0261 | · | · | · |
| W12 | 0.13× | 0.2451 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0325 0.0238 | · | · | · |
| W13 | 0.12× | 0.2387 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0377 0.0235 | · | · | · |
| W14 | 0.09× | 0.8599 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0750 0.0750 0.0750 | · | · | · |
MedianFwd¶
- W1x: [2048, 4096]dtype=bf16
hidden-state-reduce - W2x: [2048, 4096]dtype=f16
hidden-state-reduce - W3x: [64, 32768]dtype=bf16
long-seq-reduce
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.00× | 363.6075 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.5075 0.5055 | · | · | · |
| W2 | 0.00× | 363.6019 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.5104 0.5042 | · | · | · |
| W3 | 0.00× | 117.0786 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2700 0.2632 | · | · | · |
SortFwd¶
- W1x: [2048, 4096]dtype=bf16
hidden-state-reduce - W2x: [2048, 4096]dtype=f16
hidden-state-reduce - W3x: [64, 32768]dtype=bf16
long-seq-reduce
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.00× | 363.5121 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.3936 0.3886 | · | · | · |
| W2 | 0.00× | 363.5584 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.3915 0.3937 | · | · | · |
| W3 | 0.00× | 117.0429 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.2110 0.2037 | · | · | · |
DropoutFwd¶
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
Reduction¶
CummaxFwd¶
- W1x: [2048, 4096]dtype=bf16
hidden-state-scan - W2x: [2048, 4096]dtype=f16
hidden-state-scan - W3x: [64, 32768]dtype=bf16
long-seq-scan
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 5.13× | 13.8046 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 70.8431 70.8659 | · | · | · |
| W2 | 5.13× | 13.8078 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 70.8836 70.8146 | · | · | · |
| W3 | 3.52× | 5.1425 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 18.1371 17.7641 | · | · | · |
MeanVarWelfordFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-var-mean - W3x: [2048, 4096]dtype=f16
hidden-state-var-mean - W4x: [64, 32768]dtype=bf16
long-seq-var-mean
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.44× | 0.0260 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1161 0.1087 | · | · | · |
| W2 | 2.77× | 0.0530 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1437 0.1391 | · | · | · |
| W3 | 2.82× | 0.0530 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1447 0.1385 | · | · | · |
| W4 | 3.07× | 0.0271 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0798 0.0740 | · | · | · |
LogSumExpFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [32, 32, 32768]dtype=bf16
attn-weights-32k - W3x: [32, 32, 4096]dtype=bf16
attn-weights-4k - W4x: [32, 32, 4096]dtype=f16
attn-weights-4k - W5x: [4, 102400]dtype=bf16
lm-head-logits - W6x: [4, 102400]dtype=f16
lm-head-logits
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.70× | 0.0423 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5 | 0.0721 0.0660 0.0235 | · | · | · |
| W2 | 2.75× | 0.1708 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5 | 0.4632 0.4584 0.1111 | · | · | · |
| W3 | 2.77× | 0.0357 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5 | 0.1065 0.0975 0.0183 | · | · | · |
| W4 | 2.79× | 0.0352 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5 | 0.0980 0.0960 0.0168 | · | · | · |
| W5 | 2.79× | 0.0205 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5 | 0.0573 0.0515 0.0139 | · | · | · |
| W6 | 3.56× | 0.0158 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:shipped tilelang-mlir-ascend ⓘ eager · n=5 | 0.0560 0.0500 0.0144 | · | · | · |
SumFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-reduce - W3x: [2048, 4096]dtype=bf16dim=0
hidden-state-reduce-dim0 - W4x: [2048, 4096]dtype=f16
hidden-state-reduce - W5x: [2048, 4096]dtype=bf16dim=-1keepdim=true
hidden-state-reduce-keepdim - W6x: [64, 32768]dtype=bf16
long-seq-reduce
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.94× | 0.0350 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0678 0.0520 | · | · | · |
| W2 | 1.85× | 0.0372 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0688 0.0527 | · | · | · |
| W3 | 1.08× | 0.0870 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0932 0.0831 | · | · | · |
| W4 | 1.79× | 0.0375 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0686 0.0515 | · | · | · |
| W5 | 2.17× | 0.0348 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0755 0.0640 | · | · | · |
| W6 | 1.70× | 0.0205 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0344 0.0200 | · | · | · |
AminFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-reduce - W3x: [2048, 4096]dtype=f16
hidden-state-reduce - W4x: [64, 32768]dtype=bf16
long-seq-reduce
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.88× | 0.0357 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0335 0.0217 | · | · | · |
| W2 | 1.84× | 0.0372 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0651 0.0542 | · | · | · |
| W3 | 1.82× | 0.0375 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0670 0.0542 | · | · | · |
| W4 | 1.37× | 0.0230 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0333 0.0205 | · | · | · |
MeanFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-reduce - W3x: [2048, 4096]dtype=f16
hidden-state-reduce - W4x: [64, 32768]dtype=bf16
long-seq-reduce
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.77× | 0.0365 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0280 0.0565 0.0208 | · | · | · |
| W2 | 1.65× | 0.0390 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0649 0.0600 0.0535 | · | · | · |
| W3 | 1.76× | 0.0374 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0660 0.0607 0.0535 | · | · | · |
| W4 | 1.72× | 0.0198 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0340 0.0312 0.0226 | · | · | · |
AmaxFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-reduce - W3x: [2048, 4096]dtype=f16
hidden-state-reduce - W4x: [64, 32768]dtype=bf16
long-seq-reduce
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.80× | 0.0357 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0290 0.0215 | · | · | · |
| W2 | 1.85× | 0.0357 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0655 0.0515 | · | · | · |
| W3 | 1.81× | 0.0365 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0649 0.0515 | · | · | · |
| W4 | 1.41× | 0.0220 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0330 0.0192 | · | · | · |
MaximumFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill - W7input, other: [4097]dtype=bf16
nondiv-tail - W8input, other: [4097]dtype=f16
nondiv-tail - W9input, other: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.23× | 0.0626 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0651 0.0653 0.0650 | · | · | · |
| W2 | 1.06× | 0.0498 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0530 0.0678 0.0485 | · | · | · |
| W3 | 1.14× | 0.0890 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0899 0.0897 0.0899 | · | · | · |
| W4 | 1.09× | 0.0585 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0643 0.0580 0.0592 | · | · | · |
| W5 | 0.98× | 0.0573 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0851 0.0558 0.0565 0.0519 | · | · | · |
| W6 | 1.02× | 0.1026 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.1000 0.1067 0.1045 0.1022 | · | · | · |
| W7 | 1.12× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
| W8 | 0.88× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.00125 0.00125 0.00125 | · | · | · |
| W9 | 1.12× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
MinimumFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill - W7input, other: [4097]dtype=bf16
nondiv-tail - W8input, other: [4097]dtype=f16
nondiv-tail - W9input, other: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.28× | 0.0628 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0633 0.0630 0.0633 | · | · | · |
| W2 | 1.02× | 0.0515 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0519 0.0575 0.0485 | · | · | · |
| W3 | 1.12× | 0.0870 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0890 0.0891 0.0890 | · | · | · |
| W4 | 1.08× | 0.0560 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0610 0.0566 0.0580 | · | · | · |
| W5 | 0.99× | 0.0548 | tilelang-ascend tilelang-ascend ⓘtorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0849 0.0545 0.0544 0.0510 | · | · | · |
| W6 | 1.03× | 0.1003 | tilelang-ascend tilelang-ascend ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0985 0.1047 0.1032 0.1016 | · | · | · |
| W7 | 0.89× | 0.00225 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
| W8 | 0.88× | 0.002 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.00125 0.0015 | · | · | · |
| W9 | 1.00× | 0.00225 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0015 0.0015 0.0015 | · | · | · |
ArgmaxFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=0
3d-non-last-axis-argmax - W2x: [2048, 4096]dtype=bf16dim=-1
hidden-state-argmax - W3x: [2048, 4096]dtype=f16dim=-1
hidden-state-argmax - W4x: [4, 102400]dtype=bf16dim=-1
lm-head-argmax - W5x: [4, 102400]dtype=f16dim=-1
lm-head-argmax
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.08× | 0.0295 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0232 0.0232 0.0231 | · | · | · |
| W2 | 0.90× | 0.0521 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0465 0.0460 0.0466 | · | · | · |
| W3 | 0.75× | 0.0522 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0390 0.0380 0.0387 | · | · | · |
| W4 | 1.01× | 0.0174 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0175 0.0213 0.0173 | · | · | · |
| W5 | 0.59× | 0.0175 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0103 0.0145 0.00975 | · | · | · |
ArgminFwd¶
dim=-1
- W1x: [2048, 4096]dtype=bf16
hidden-state-argmin - W2x: [2048, 4096]dtype=f16
hidden-state-argmin - W3x: [4, 102400]dtype=bf16
lm-head-argmin - W4x: [4, 102400]dtype=f16
lm-head-argmin
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.93× | 0.0542 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0503 0.0460 0.0454 | · | · | · |
| W2 | 0.70× | 0.0540 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0385 0.0376 0.0367 | · | · | · |
| W3 | 0.86× | 0.0203 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0175 0.0213 0.0163 | · | · | · |
| W4 | 0.49× | 0.0187 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.00925 0.0138 0.0085 | · | · | · |
SegmentSumFwd¶
- W1x: [2048, 4096]dtype=bf16
hidden-state-reduce - W2x: [2048, 4096]dtype=f16
hidden-state-reduce - W3x: [64, 32768]dtype=bf16
long-seq-reduce
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.27× | 0.2995 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0809 0.0775 | · | · | · |
| W2 | 0.28× | 0.3003 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0875 0.0780 | · | · | · |
| W3 | 0.47× | 0.0975 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0457 0.0400 | · | · | · |
LogSoftmaxFwd¶
- W1x: [32, 32, 32768]dtype=bf16
attn-weights-32k - W2x: [32, 32, 4096]dtype=bf16
attn-weights-4k - W3x: [32, 32, 4096]dtype=f16
attn-weights-4k - W4x: [32, 32, 4096]dtype=f32
attn-weights-4k - W5x: [4, 102400]dtype=bf16
lm-head-logits - W6x: [4, 102400]dtype=f16
lm-head-logits - W7x: [4, 102400]dtype=f32
lm-head-logits
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.63× | 0.8110 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.4966 0.4846 0.4911 | · | · | · |
| W2 | 0.65× | 0.1275 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0850 0.0665 0.0664 | · | · | · |
| W3 | 0.67× | 0.1231 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0815 0.0668 0.0660 | · | · | · |
| W4 | 0.36× | 0.1618 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0583 0.0540 0.0460 | · | · | · |
| W5 | 0.16× | 0.3415 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0408 0.0406 0.0408 | · | · | · |
| W6 | 0.16× | 0.3337 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0405 0.0408 0.0405 | · | · | · |
| W7 | 0.10× | 0.3897 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0374 0.0375 0.0372 | · | · | · |
SoftmaxFwd¶
- W1x: [32, 32, 32768]dtype=bf16
attn-weights-32k - W2x: [32, 32, 4096]dtype=bf16
attn-weights-4k - W3x: [32, 32, 4096]dtype=f16
attn-weights-4k - W4x: [32, 32, 4096]dtype=f32
attn-weights-4k - W5x: [4, 102400]dtype=bf16
lm-head-logits - W6x: [4, 102400]dtype=f16
lm-head-logits - W7x: [4, 102400]dtype=f32
lm-head-logits
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.57× | 0.8177 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.4669 0.4818 0.4929 | · | · | · |
| W2 | 0.55× | 0.1283 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0727 0.0630 0.0658 | · | · | · |
| W3 | 0.55× | 0.1269 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0695 0.0641 0.0653 | · | · | · |
| W4 | 0.31× | 0.1613 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0495 0.0532 0.0474 | · | · | · |
| W5 | 0.15× | 0.3559 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0425 0.0423 0.0423 | · | · | · |
| W6 | 0.15× | 0.3548 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0423 0.0423 0.0419 | · | · | · |
| W7 | 0.09× | 0.4040 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor librarytilelang-mlir-ascend:autotune_2d tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0380 0.0381 0.0382 0.1042 | · | · | · |
MaskedReduceSumFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-reduce - W3x: [2048, 4096]dtype=bf16dim=0
hidden-state-reduce-dim0 - W4x: [2048, 4096]dtype=f16
hidden-state-reduce - W5x: [2048, 4096]dtype=bf16dim=-1keepdim=true
hidden-state-reduce-keepdim - W6x: [64, 32768]dtype=bf16
long-seq-reduce
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.16× | 0.6412 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0995 0.0848 | · | · | · |
| W2 | 0.07× | 2.3206 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1578 0.1417 | · | · | · |
| W3 | 0.08× | 2.3537 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1747 0.1653 | · | · | · |
| W4 | 0.07× | 2.3205 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1562 0.1405 | · | · | · |
| W5 | 0.08× | 2.3161 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1570 0.1469 | · | · | · |
| W6 | 0.11× | 0.6468 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0737 0.0624 | · | · | · |