Second-phase operators¶
50 ops, 236 workloads — Attention 9 · Linear Attention / SSM 2 · Mixture of Experts 1 · Quantization 3 · Elementwise 22 · Reduction 1 · Normalization 2 · Pooling 2 · Positional Encoding 6 · MHC 2.
One table per op, one row per workload. Ratio is the baseline device time divided by ours in the same measurement regime, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.
A separate set, and a separate denominator
These operators are not in oplist150, the first-phase acceptance list (docs/reports/R267-data/oplist150.json) that the overview and the per-family data pages are scoped to and that progress is reported against. They are shown here because the harness already measured a real opponent against them, not because the first phase grew.
Nothing on this page is added to any count on the other pages. The overview's coverage figures, the per-family tallies and the attainment counts are all first-phase only, and stay first-phase only. This page's own tally, below, counts these operators and nothing else.
An operator reaches this page only when the snapshot publishes a ratio for it. One with no timed opponent is left out rather than shown as a row of blanks.
Attention¶
GroupedQueryAttentionDecodePagedWithKVCacheFwd¶
dtype=f16
- W1q: [4, 128, 128]kv: [4096, 8, 128]page_size=256
serving-405b-p256 - W2q: [8, 64, 128]kv: [4096, 8, 128]page_size=256
serving-70b-p256 - W3q: [8, 64, 128]kv: [4096, 8, 128]page_size=64
serving-70b-p64 - W4q: [8, 32, 128]kv: [32768, 8, 128]page_size=64
serving-8b-long-p64 - W5q: [32, 32, 128]kv: [4096, 8, 128]page_size=256
serving-8b-p256 - W6q: [32, 32, 128]kv: [4096, 8, 128]page_size=64
serving-8b-p64 - W7q: [64, 32, 128]kv: [2048, 8, 128]page_size=64
throughput-8b-p64
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.45× | 33.4937 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager python reference vendor library | 48.4476 48.5021 48.5410 | · | · | · |
| W2 | 1.19× | 33.4759 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager python reference vendor library | 39.9049 39.9079 44.9059 | · | · | · |
| W3 | — | · | torch_compile_ge torch_compile basistorch_npu eager python reference vendor library | 39.6264 44.8341 | · | · | · |
| W4 | — | · | torch_compile_ge torch_compile basistorch_npu eager python reference vendor library | 136.3076 165.6499 | · | · | · |
| W5 | 1.18× | 65.3830 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager python reference vendor library | 77.1264 77.1175 87.2837 | · | · | · |
| W6 | — | · | torch_compile_ge torch_compile basistorch_npu eager python reference vendor library | 76.0240 88.2759 | · | · | · |
| W7 | — | · | torch_compile_ge torch_compile basistorch_npu eager python reference vendor library | 144.2580 161.4332 | · | · | · |
GroupedQueryAttentionSlidingWindowVarlenFwd¶
dtype=f16batch=4dim=128
- W1total_q=8192total_k=8192heads=64/8window_size_left=1024max_seqlen_q=2048
llama-70b-long-w1024 - W2total_q=2048total_k=2048heads=64/8window_size_left=256max_seqlen_q=512
llama-70b-short-w256 - W3total_q=8192total_k=8192heads=32/8window_size_left=1024max_seqlen_q=2048
llama-8b-long-w1024 - W4total_q=2048total_k=2048heads=32/8window_size_left=256max_seqlen_q=512
llama-8b-short-w256
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.81× | 6.2363 | torch_compile_aclgraph torch_compiletorch_npu eager python reference vendor library basis | 5.0464 5.0319 | · | · | · |
| W2 | 1.16× | 0.5373 | torch_compile_aclgraph torch_compile basistorch_npu eager python reference vendor library | 0.6249 0.6279 | · | · | · |
| W3 | 0.80× | 3.3505 | torch_compile_aclgraph torch_compile basistorch_npu eager python reference vendor library | 2.6565 2.7251 | · | · | · |
| W4 | 1.26× | 0.3385 | torch_compile_aclgraph torch_compiletorch_npu eager python reference vendor library basis | 0.4236 0.4211 | · | · | · |
MultiHeadAttentionFwd¶
- W1q, kv: [1, 2048, 64, 128]dtype=bf16
llama-70b-long - W2q, kv: [1, 2048, 64, 128]dtype=f16
llama-70b-long - W3q, kv: [2, 512, 64, 128]dtype=bf16
llama-70b-short - W4q, kv: [2, 512, 64, 128]dtype=f16
llama-70b-short - W5q, kv: [2, 2048, 32, 128]dtype=bf16
llama-8b-long - W6q, kv: [2, 2048, 32, 128]dtype=f16
llama-8b-long - W7q, kv: [4, 512, 32, 128]dtype=bf16
llama-8b-short - W8q, kv: [4, 512, 32, 128]dtype=f16
llama-8b-short
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.42× | 1.5755 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager SDPA vendor library | 0.6620 0.8296 0.8330 0.8325 | · | · | · |
| W2 | 0.42× | 1.5684 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.6521 0.8214 0.8235 | · | · | · |
| W3 | 0.63× | 0.2579 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.1611 0.2915 0.2884 | · | · | · |
| W4 | 0.65× | 0.2534 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.1608 0.2904 0.2855 | · | · | · |
| W5 | 0.42× | 1.5901 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.6626 0.8350 0.8306 | · | · | · |
| W6 | 0.41× | 1.5860 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.6498 0.8173 0.8165 | · | · | · |
| W7 | 0.64× | 0.2560 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.1616 0.2949 0.2945 | · | · | · |
| W8 | 0.66× | 0.2555 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.1638 0.2984 0.2961 | · | · | · |
GroupedQueryAttentionFwd¶
dtype=f16
- W1q: [1, 2048, 64, 128]kv: [1, 2048, 8, 128]
llama-70b-long - W2q: [2, 512, 64, 128]kv: [2, 512, 8, 128]
llama-70b-short - W3q: [2, 2048, 32, 128]kv: [2, 2048, 8, 128]
llama-8b-long - W4q: [4, 512, 32, 128]kv: [4, 512, 8, 128]
llama-8b-short
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.41× | 1.5300 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.6172 0.7180 0.7185 | · | · | · |
| W2 | 0.60× | 0.2390 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.1411 0.2612 0.2616 | · | · | · |
| W3 | 0.41× | 1.5433 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.6211 0.7411 0.7349 | · | · | · |
| W4 | 0.60× | 0.2439 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library | 0.1449 0.2695 0.2679 | · | · | · |
GroupedQueryAttentionSlidingWindowFwd¶
dtype=f16
- W1q: [1, 2048, 64, 128]kv: [1, 2048, 8, 128]window_size_left=1024
llama-70b-long-w1024 - W2q: [2, 512, 64, 128]kv: [2, 512, 8, 128]window_size_left=256
llama-70b-short-w256 - W3q: [2, 2048, 32, 128]kv: [2, 2048, 8, 128]window_size_left=1024
llama-8b-long-w1024 - W4q: [4, 512, 32, 128]kv: [4, 512, 8, 128]window_size_left=256
llama-8b-short-w256
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.30× | 1.8444 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager python reference vendor library | 0.5539 1.5036 1.5085 | · | · | · |
| W2 | 0.43× | 0.3362 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager python reference vendor library | 0.1391 0.3902 0.3925 | · | · | · |
| W3 | 0.30× | 1.8560 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager python reference vendor library | 0.5573 1.5281 1.5203 | · | · | · |
| W4 | 0.43× | 0.3429 | ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager python reference vendor library | 0.1434 0.3946 0.3921 | · | · | · |
MultiHeadAttentionDecodeWithKVCacheFwd¶
- W1q: [4, 1, 64, 128]kv: [4, 32768, 64, 128]dtype=bf16
llama-70b-32k - W2q: [4, 1, 64, 128]kv: [4, 32768, 64, 128]dtype=f16
llama-70b-32k - W3q: [16, 1, 64, 128]kv: [16, 4096, 64, 128]dtype=bf16
llama-70b-4k - W4q: [16, 1, 64, 128]kv: [16, 4096, 64, 128]dtype=f16
llama-70b-4k - W5q: [8, 1, 32, 128]kv: [8, 32768, 32, 128]dtype=bf16
llama-8b-32k - W6q: [8, 1, 32, 128]kv: [8, 32768, 32, 128]dtype=f16
llama-8b-32k - W7q: [32, 1, 32, 128]kv: [32, 4096, 32, 128]dtype=bf16
llama-8b-4k - W8q: [32, 1, 32, 128]kv: [32, 4096, 32, 128]dtype=f16
llama-8b-4k
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.13× | 139.8074 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager SDPA vendor library basis | 18.6446 19.0927 18.4155 | · | · | · |
| W2 | 0.13× | 139.8650 | torch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library basis | 18.7369 18.6160 | · | · | · |
| W3 | 0.14× | 65.0734 | torch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library basis | 9.2641 9.2630 | · | · | · |
| W4 | 0.14× | 65.1240 | torch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library basis | 9.2516 9.2351 | · | · | · |
| W5 | 0.15× | 139.7050 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA vendor library | 20.7512 20.7583 | · | · | · |
| W6 | 0.15× | 139.7111 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA vendor library | 20.6946 20.7183 | · | · | · |
| W7 | 0.14× | 65.0439 | torch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library basis | 9.0760 9.0411 | · | · | · |
| W8 | 0.14× | 65.1153 | torch_compile_aclgraph torch_compiletorch_npu eager SDPA vendor library basis | 9.0676 9.0475 | · | · | · |
GroupedQueryAttentionBwd¶
dtype=f16
- W1q: [1, 2048, 64, 128]kv: [1, 2048, 8, 128]
llama-70b-long - W2q: [2, 512, 64, 128]kv: [2, 512, 8, 128]
llama-70b-short - W3q: [2, 2048, 32, 128]kv: [2, 2048, 8, 128]
llama-8b-long - W4q: [4, 512, 32, 128]kv: [4, 512, 8, 128]
llama-8b-short
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.01× | 270.8478 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA autograd vendor library | 2.3700 2.3945 | · | · | · |
| W2 | 0.02× | 44.8023 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA autograd vendor library | 0.7984 0.8086 | · | · | · |
| W3 | 0.01× | 272.6906 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA autograd vendor library | 2.4531 2.4635 | · | · | · |
| W4 | 0.02× | 45.1514 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA autograd vendor library | 0.7960 0.8096 | · | · | · |
MultiHeadAttentionBwd¶
dtype=f16
- W1q, kv: [1, 2048, 64, 128]
llama-70b-long - W2q, kv: [2, 512, 64, 128]
llama-70b-short - W3q, kv: [2, 2048, 32, 128]
llama-8b-long - W4q, kv: [4, 512, 32, 128]
llama-8b-short
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.01× | 365.0621 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA autograd vendor library | 2.9143 2.9274 | · | · | · |
| W2 | 0.01× | 91.6334 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA autograd vendor library | 0.7804 0.7946 | · | · | · |
| W3 | 0.01× | 365.5979 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA autograd vendor library | 2.8984 2.9120 | · | · | · |
| W4 | 0.01× | 91.8776 | torch_compile_aclgraph torch_compile basistorch_npu eager SDPA autograd vendor library | 0.7937 0.7959 | · | · | · |
MultiHeadLatentAttentionDecodeWithKVCacheFwd¶
dtype=f16pe_dim=64
- W1q: [8, 128, 128]q_pe: [8, 128, 64]kv: [8, 32768, 1, 128]k_pe: [8, 32768, 1, 64]
deepseek-v2-32k - W2q: [32, 128, 128]q_pe: [32, 128, 64]kv: [32, 4096, 1, 128]k_pe: [32, 4096, 1, 64]
deepseek-v2-4k
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.00× | 521.9301 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager python reference vendor library basis | 0.6825 0.6180 0.6683 | · | · | · |
| W2 | 0.00× | 260.9430 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager python reference vendor library basis | 0.3005 0.2662 0.2871 | · | · | · |
Linear Attention / SSM¶
EngramGateConvBwd¶
dY, H, k, v, vhat: [M, seq_len, d]rms_w_h, rms_w_v: [d]conv_w: [4, d]alpha, rrms_h, rrms_k, rrms_v: [M, seq_len], f32
- W1dtype=bf16M=1seq_len=128d=256
bwd-b1-s128-d256 - W2dtype=f16M=1seq_len=32d=256
bwd-b1-s32-d256 - W3dtype=f16M=2seq_len=64d=512
bwd-b2-s64-d512
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 11.01× | 0.0831 | Torch-NPU eager exact reference | 0.9155 | · | · | · |
| W2 | 17.36× | 0.0408 | Torch-NPU eager exact reference | 0.7074 | · | · | · |
| W3 | 12.66× | 0.0825 | Torch-NPU eager exact reference | 1.0442 | · | · | · |
EngramGateConvFwd¶
H, k, v: [M, seq_len, d]rms_w_h, rms_w_v: [d]conv_w: [4, d]
- W1dtype=bf16M=1seq_len=128d=256
fwd-b1-s128-d256 - W2dtype=f16M=1seq_len=32d=256
fwd-b1-s32-d256 - W3dtype=f16M=2seq_len=64d=512
fwd-b2-s64-d512
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 8.71× | 0.0410 | Torch-NPU eager exact reference | 0.3571 | · | · | · |
| W2 | 9.95× | 0.0280 | Torch-NPU eager exact reference | 0.2785 | · | · | · |
| W3 | 8.76× | 0.0452 | Torch-NPU eager exact reference | 0.3962 | · | · | · |
Mixture of Experts¶
MoeUnpermuteFwd¶
dtype=bf16top_k=8
- W1total_tokens=1hidden_size=7168
large-hidden-decode - W2total_tokens=512hidden_size=7168
large-hidden-medium - W3total_tokens=4096hidden_size=7168
large-hidden-prefill - W4total_tokens=32hidden_size=7168
large-hidden-small - W5total_tokens=1hidden_size=3072
small-hidden-decode - W6total_tokens=512hidden_size=3072
small-hidden-medium - W7total_tokens=4096hidden_size=3072
small-hidden-prefill - W8total_tokens=32hidden_size=3072
small-hidden-small
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.84× | 0.00625 | ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.00425 0.0445 0.0245 | · | · | · |
| W2 | 0.51× | 0.1747 | ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.0870 0.3869 0.3757 | · | · | · |
| W3 | 0.41× | 1.3121 | ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.5310 2.8803 2.8481 | · | · | · |
| W4 | 0.64× | 0.0180 | ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.00975 0.0600 0.0454 | · | · | · |
| W5 | 0.90× | 0.005 | ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.00313 0.0380 0.0180 | · | · | · |
| W6 | 0.35× | 0.1300 | ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.0420 0.1510 0.1355 | · | · | · |
| W7 | 0.29× | 0.9976 | ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.2871 1.2434 1.2379 | · | · | · |
| W8 | 0.64× | 0.0152 | ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.00675 0.0561 0.0400 | · | · | · |
Quantization¶
FP8LightningIndexerFwd¶
index_q: [batch, seq_len, heads, index_dim]index_k: [batch, seq_len_kv, kv_group, index_dim]weights: [seq_len, heads], f32cu_seqlen_ks, cu_seqlen_ke: [seq_len], i32index_k_scale: [batch, seq_len_kv, kv_group], f32
batch=1seq_len=8192heads=32index_dim=64seq_len_kv=32768kv_group=1
- W1dtype=bf16
lightning-indexer-s8k-h32-d64 - W2dtype=fp8e4m3
lightning-indexer-s8k-h32-d64
index_k_scale: [batch, seq_len_kv, kv_group], f32
dtype=fp8e4m3batch=1seq_len_kv=32768kv_group=1seq_len=8192heads=32index_dim=64
- W3
lightning-indexer-s8k-h32-d64-kscale
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 5.88× | 24.0535 | torch_npu eager python reference vendor library basis ⓘ | 141.4828 | · | · | · |
| W2 | 5.82× | 24.2429 | torch_npu eager python reference vendor library basis ⓘ | 141.1920 | · | · | · |
| W3 | 5.70× | 25.1346 | torch_npu eager python reference vendor library basis ⓘ | 143.3777 | · | · | · |
BmmFp8NKFwd¶
dtype=fp8e4m3
- W1b=32m=128n=128k=2048
mha-decode-b32-pv-per-tensor - W2b=64m=128n=2048k=128
mha-decode-b64-qk-per-tensor - W3b=128m=512n=512k=2048
moe-prefill-b128-per-tensor - W4b=4m=1024n=1024k=1024
square-b4-1k-per-tensor - W5b=8m=2048n=2048k=2048
square-b8-2k-per-tensor
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.83× | 0.1133 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.5479 0.5954 | · | · | · |
| W2 | 3.65× | 0.2238 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.8139 0.9902 | · | · | · |
| W3 | 9.28× | 2.3949 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 22.3016 22.2772 | · | · | · |
| W4 | 2.42× | 0.1383 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.3374 0.3259 | · | · | · |
| W5 | 3.48× | 1.4383 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 5.0110 5.5925 | · | · | · |
BmmFp8KNFwd¶
dtype=fp8e4m3
- W1b=32m=128n=128k=2048
mha-decode-b32-pv-per-tensor - W2b=64m=128n=2048k=128
mha-decode-b64-qk-per-tensor - W3b=128m=512n=512k=2048
moe-prefill-b128-per-tensor - W4b=4m=1024n=1024k=1024
square-b4-1k-per-tensor - W5b=8m=2048n=2048k=2048
square-b8-2k-per-tensor
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.12× | 0.1305 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.5369 0.5880 | · | · | · |
| W2 | 3.87× | 0.2070 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 0.8036 0.9409 | · | · | · |
| W3 | 9.28× | 2.3847 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 22.2211 22.1105 | · | · | · |
| W4 | 2.36× | 0.1400 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.3287 0.3186 | · | · | · |
| W5 | 3.50× | 1.4418 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 5.0496 5.6059 | · | · | · |
Elementwise¶
ClampScalarFwd¶
min=-0.5max=0.5
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.18× | 0.0691 | ClipByValueV2_e87c9546cf2804355cf0d9445a2db01a_high_performance_220010000 | 0.2890 | · | · | · |
| W2 | 2.23× | 0.0683 | ClipByValueV2_1056052785b43db3ca4234fbf5e08c0e_high_performance_220010000 | 0.1522 | · | · | · |
| W3 | 2.50× | 0.1150 | ClipByValueV2_54f1d1caf8fe36c496f699fefe98fbca_high_performance_220010000 | 0.2878 | · | · | · |
| W4 | 4.73× | 0.9459 | ClipByValueV2_e87c9546cf2804355cf0d9445a2db01a_high_performance_220010000 | 4.4703 | · | · | · |
| W5 | 2.35× | 0.9565 | ClipByValueV2_1056052785b43db3ca4234fbf5e08c0e_high_performance_220010000 | 2.2517 | · | · | · |
BitwiseNotFwd¶
- W1input: [4096, 4096]dtype=i32
elementwise-16M - W2input: [4096, 4096]dtype=i64
elementwise-16M - W3input: [16384, 16384]dtype=i32
elementwise-256M - W4input: [4097]dtype=i32
legacy-probe-1 - W5input: [4097]dtype=i64
legacy-probe-2
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.53× | 5.0855 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 22.9669 22.9959 | · | · | · |
| W2 | 6.08× | 5.1049 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 31.4594 31.4969 | · | · | · |
| W3 | 4.53× | 80.5434 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library ⓘ | 364.8630 364.9903 | · | · | · |
| W4 | 0.77× | 0.0410 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0325 0.0284 | · | · | · |
| W5 | 0.83× | 0.0431 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0333 0.0297 | · | · | · |
AlibiFwd¶
num_heads=32
- W1dtype=bf16seq_len=2048
llama-prefill-2k - W2dtype=f16seq_len=2048
llama-prefill-2k - W3dtype=bf16seq_len=4096
llama-prefill-4k - W4dtype=f16seq_len=4096
llama-prefill-4k
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.25× | 0.5493 | Torch-NPU eager reference | 1.2385 | · | · | · |
| W2 | 2.25× | 0.5505 | Torch-NPU eager reference | 1.2361 | · | · | · |
| W3 | 2.16× | 2.1581 | Torch-NPU eager reference | 4.6643 | · | · | · |
| W4 | 2.16× | 2.1609 | Torch-NPU eager reference | 4.6753 | · | · | · |
PreluFwd¶
- W1input: [1, 4, 1025]weight: [4]dtype=f16
channel-tail-nondivisible - W2input: [16, 256, 56, 56]weight: [256]dtype=bf16
cnn-feat-per-channel - W3input: [16, 512, 28, 28]weight: [512]dtype=bf16
cnn-feat-per-channel-deep - W4input: [16, 512, 28, 28]weight: [512]dtype=f16
cnn-feat-per-channel-deep - W5input: [16, 256, 56, 56]weight: [256]dtype=f16
cnn-feat-per-channel
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.61× | 0.0248 | Torch-NPU eager exact reference | 0.0150 | · | · | · |
| W2 | 2.43× | 0.0620 | Torch-NPU eager exact reference | 0.1509 | · | · | · |
| W3 | 2.41× | 0.0379 | Torch-NPU eager exact reference | 0.0912 | · | · | · |
| W4 | 2.41× | 0.0380 | Torch-NPU eager exact reference | 0.0917 | · | · | · |
| W5 | 2.30× | 0.0670 | Torch-NPU eager exact reference | 0.1541 | · | · | · |
MaskedFillScalarFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.67× | 0.0843 | Torch-NPU eager exact reference | 0.1410 | · | · | · |
| W2 | 1.25× | 0.0848 | Torch-NPU eager exact reference | 0.1060 | · | · | · |
| W3 | 1.38× | 0.1395 | Torch-NPU eager exact reference | 0.1925 | · | · | · |
| W4 | 2.18× | 1.1220 | Torch-NPU eager exact reference | 2.4454 | · | · | · |
| W5 | 1.85× | 1.1454 | Torch-NPU eager exact reference | 2.1162 | · | · | · |
LogicalOrFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bool
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W4input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W5input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=bool
hidden-state-prefill - W7input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W8input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.73× | 0.0612 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0998 0.0998 0.0998 | · | · | · |
| W2 | 0.94× | 0.0452 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library | 0.0359 0.0356 0.0360 | · | · | · |
| W3 | 1.83× | 0.0512 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0835 0.0833 0.0835 | · | · | · |
| W4 | 1.60× | 0.0895 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1360 0.1363 0.1360 | · | · | · |
| W5 | 1.71× | 0.0576 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0990 0.0931 0.0830 | · | · | · |
| W6 | 1.00× | 0.0356 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0355 0.0325 0.0297 | · | · | · |
| W7 | 1.53× | 0.0559 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0828 0.0770 0.0674 | · | · | · |
| W8 | 1.37× | 0.1034 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1415 0.1325 0.1224 | · | · | · |
NeFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.81× | 0.0658 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.1077 0.1077 0.1077 | · | · | · |
| W2 | 1.01× | 0.0563 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0560 0.0548 0.0530 | · | · | · |
| W3 | 1.29× | 0.0945 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.1163 0.1160 0.1163 | · | · | · |
| W4 | 1.37× | 0.0644 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0885 0.0794 0.0793 | · | · | · |
| W5 | 0.96× | 0.0616 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0592 0.0607 0.0553 | · | · | · |
| W6 | 1.16× | 0.1062 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.1217 0.1123 0.1153 | · | · | · |
GtFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.22× | 0.0595 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0653 0.0650 0.0651 | · | · | · |
| W2 | 1.06× | 0.0512 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0540 0.0678 0.0464 | · | · | · |
| W3 | 1.10× | 0.0835 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0833 0.0833 0.0831 | · | · | · |
| W4 | 1.15× | 0.0545 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0626 0.0550 0.0548 | · | · | · |
| W5 | 0.93× | 0.0568 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0530 0.0510 0.0485 | · | · | · |
| W6 | 1.04× | 0.0943 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0980 0.0998 0.0935 | · | · | · |
LtFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.21× | 0.0578 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0602 0.0602 0.0600 | · | · | · |
| W2 | 0.98× | 0.0455 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0480 0.0633 0.0415 | · | · | · |
| W3 | 1.18× | 0.0789 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0789 0.0788 0.0786 | · | · | · |
| W4 | 1.16× | 0.0504 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0580 0.0532 0.0493 | · | · | · |
| W5 | 0.95× | 0.0478 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0451 0.0430 0.0393 | · | · | · |
| W6 | 1.01× | 0.0917 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0960 0.0944 0.0902 | · | · | · |
LeFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.16× | 0.0553 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.0573 0.0575 0.0574 | · | · | · |
| W2 | 1.01× | 0.0457 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0460 0.0558 0.0405 | · | · | · |
| W3 | 1.17× | 0.0756 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0747 0.0751 0.0746 | · | · | · |
| W4 | 1.15× | 0.0483 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0553 0.0503 0.0480 | · | · | · |
| W5 | 0.96× | 0.0483 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0465 0.0455 0.0428 | · | · | · |
| W6 | 1.03× | 0.0880 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0877 0.0876 0.0846 | · | · | · |
GeFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.16× | 0.0633 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.0673 0.0673 0.0674 | · | · | · |
| W2 | 1.10× | 0.0452 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0498 0.0658 0.0430 | · | · | · |
| W3 | 1.08× | 0.0830 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0833 0.0835 0.0833 | · | · | · |
| W4 | 1.10× | 0.0537 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0575 0.0537 0.0514 | · | · | · |
| W5 | 0.96× | 0.0505 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0498 0.0491 0.0447 | · | · | · |
| W6 | 1.04× | 0.0958 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.1000 0.1010 0.0970 | · | · | · |
EqFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.13× | 0.0665 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0653 0.0653 0.0651 | · | · | · |
| W2 | 1.03× | 0.0560 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0573 0.0575 0.0527 | · | · | · |
| W3 | 1.09× | 0.0932 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0927 0.0927 0.0925 | · | · | · |
| W4 | 1.10× | 0.0633 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0700 0.0633 0.0636 | · | · | · |
| W5 | 0.94× | 0.0620 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0587 0.0580 0.0550 | · | · | · |
| W6 | 1.04× | 0.1057 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.1082 0.1065 0.1045 | · | · | · |
NanToNumFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.11× | 0.0798 | NanToNum_71a4100bd8f51c70f591051c16121929_high_performance_210010000 | 0.0885 | · | · | · |
| W2 | 0.86× | 0.0786 | NanToNum_ccb78600df288a689f3cc4abbf7e007e_high_performance_210010000 | 0.0673 | · | · | · |
| W3 | 0.95× | 0.1295 | NanToNum_5703d6e339578b11bc6d39a70205aa62_high_performance_210010000 | 0.1229 | · | · | · |
| W4 | 1.08× | 1.0601 | NanToNum_71a4100bd8f51c70f591051c16121929_high_performance_210010000 | 1.1493 | · | · | · |
| W5 | 0.88× | 1.0551 | NanToNum_ccb78600df288a689f3cc4abbf7e007e_high_performance_210010000 | 0.9250 | · | · | · |
SinusoidalFwd¶
d_model=4096
- W1dtype=bf16seq_len=2048
transformer-2k-4k - W2dtype=f16seq_len=2048
transformer-2k-4k - W3dtype=bf16seq_len=4096
transformer-4k-4k - W4dtype=f16seq_len=4096
transformer-4k-4k
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.02× | 0.4160 | Torch-NPU eager reference | 0.4228 | · | · | · |
| W2 | 1.03× | 0.4193 | Torch-NPU eager reference | 0.4328 | · | · | · |
| W3 | 0.91× | 0.8217 | Torch-NPU eager reference | 0.7468 | · | · | · |
| W4 | 0.91× | 0.8220 | Torch-NPU eager reference | 0.7458 | · | · | · |
SeluFwd¶
- W1input: [2048, 4096]dtype=bf16
snn-fc - W2input: [2048, 4096]dtype=f16
snn-fc - W3input: [2048, 8192]dtype=bf16
snn-fc-wide - W4input: [2048, 8192]dtype=f16
snn-fc-wide
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.84× | 0.0440 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0367 0.0367 0.0318 | · | · | · |
| W2 | 0.78× | 0.0460 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0360 0.0376 0.0323 0.0296 | · | · | · |
| W3 | 0.84× | 0.0779 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.0663 0.0661 0.0612 | · | · | · |
| W4 | 0.84× | 0.0770 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0643 0.0648 0.0607 0.0586 | · | · | · |
FloorDivideFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bf16
hidden-state-prefill - W5input, other: [2048, 4096]dtype=f16
hidden-state-prefill - W6input, other: [2048, 4096]dtype=f32
hidden-state-prefill - W7input, other: [4097]dtype=bf16
nondiv-tail - W8input, other: [4097]dtype=f16
nondiv-tail - W9input, other: [4097]dtype=f32
nondiv-tail
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.52× | 0.1578 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0814 0.0650 | · | · | · |
| W2 | 0.44× | 0.1555 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0688 0.0586 | · | · | · |
| W3 | 0.61× | 0.1580 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0993 0.0917 | · | · | · |
| W4 | 0.61× | 0.1052 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0620 0.0563 | · | · | · |
| W5 | 0.60× | 0.1054 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0639 0.0573 | · | · | · |
| W6 | 0.87× | 0.1227 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1062 0.1015 | · | · | · |
| W7 | 0.69× | 0.00325 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.00225 0.0015 | · | · | · |
| W8 | 0.75× | 0.003 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.00225 0.0015 | · | · | · |
| W9 | 0.75× | 0.003 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0015 0.0015 | · | · | · |
BitwiseXorFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bool
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i32
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i64
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bool
hidden-state-prefill - W5input, other: [2048, 4096]dtype=i32
hidden-state-prefill - W6input, other: [2048, 4096]dtype=i64
hidden-state-prefill - W7input, other: [4097]dtype=i32
tail-nondivisible
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.01× | 4.3217 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0602 0.0455 | · | · | · |
| W2 | 0.02× | 7.0530 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1075 0.0986 | · | · | · |
| W3 | 0.02× | 10.1686 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1926 0.1872 | · | · | · |
| W4 | 1.15× | 0.0333 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0379 0.0306 | · | · | · |
| W5 | 1.04× | 0.1020 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1062 0.1027 | · | · | · |
| W6 | 1.07× | 0.1956 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.2102 0.2025 | · | · | · |
| W7 | 0.89× | 0.00225 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.002 0.0015 | · | · | · |
BitwiseAndFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bool
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i32
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i64
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bool
hidden-state-prefill - W5input, other: [2048, 4096]dtype=i32
hidden-state-prefill - W6input, other: [2048, 4096]dtype=i64
hidden-state-prefill - W7input, other: [4097]dtype=i32
tail-nondivisible
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.01× | 4.3252 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0483 0.0350 | · | · | · |
| W2 | 0.01× | 7.0652 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0993 0.0912 | · | · | · |
| W3 | 0.02× | 10.1616 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1905 0.1815 | · | · | · |
| W4 | 1.08× | 0.0328 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0362 0.0312 | · | · | · |
| W5 | 1.00× | 0.1051 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1057 0.1025 | · | · | · |
| W6 | 1.04× | 0.1993 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.2055 0.2004 | · | · | · |
| W7 | 1.06× | 0.002 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.00225 0.0015 | · | · | · |
BitwiseOrFwd¶
- W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bool
cnn-feat-broadcast - W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i32
cnn-feat-broadcast - W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i64
cnn-feat-broadcast
- W4input, other: [2048, 4096]dtype=bool
hidden-state-prefill - W5input, other: [2048, 4096]dtype=i32
hidden-state-prefill - W6input, other: [2048, 4096]dtype=i64
hidden-state-prefill - W7input, other: [4097]dtype=i32
tail-nondivisible
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.01× | 4.3249 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0504 0.0357 | · | · | · |
| W2 | 0.01× | 7.0569 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0998 0.0920 | · | · | · |
| W3 | 0.02× | 10.1641 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1883 0.1815 | · | · | · |
| W4 | 1.07× | 0.0318 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0339 0.0296 | · | · | · |
| W5 | 1.01× | 0.1040 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1070 0.1010 | · | · | · |
| W6 | 1.04× | 0.1968 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.2065 0.1980 | · | · | · |
| W7 | 1.00× | 0.002 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.00225 0.0015 | · | · | · |
IsinfFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.02× | 5.9417 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0930 0.0571 0.0809 | · | · | · |
| W2 | 0.01× | 6.0650 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0770 0.0555 0.0678 | · | · | · |
| W3 | 0.02× | 5.9460 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.1360 0.1035 0.1303 | · | · | · |
| W4 | 0.02× | 94.2352 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 1.7075 1.7066 1.7060 | · | · | · |
| W5 | 0.02× | 94.1644 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 1.6007 0.7289 1.5947 | · | · | · |
IsnanFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.02× | 6.5216 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1148 0.1067 | · | · | · |
| W2 | 0.01× | 8.8604 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0587 0.0542 | · | · | · |
| W3 | 0.02× | 6.5269 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1309 0.1212 | · | · | · |
| W4 | 0.02× | 103.4591 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 1.6560 1.6482 | · | · | · |
| W5 | 0.01× | 103.3778 | torch_compile_aclgraph torch_compile basistorch_npu eager vendor library | 0.7666 0.7720 | · | · | · |
IsfiniteFwd¶
- W1input: [4096, 4096]dtype=bf16
elementwise-16M - W2input: [4096, 4096]dtype=f16
elementwise-16M - W3input: [4096, 4096]dtype=f32
elementwise-16M - W4input: [16384, 16384]dtype=bf16
elementwise-256M - W5input: [16384, 16384]dtype=f16
elementwise-256M
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.01× | 5.9420 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.2279 7.7950 0.0555 | · | · | · |
| W2 | 0.01× | 6.0667 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.1160 7.7988 0.0545 | · | · | · |
| W3 | 0.02× | 5.9465 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.2697 7.8072 0.0985 | · | · | · |
| W4 | 0.01× | 94.2362 | torch_compile_aclgraph torch_compile basistorch_compile_ge torch_compiletorch_npu eager vendor library ⓘ | 0.7360 0.7370 0.7366 | · | · | · |
| W5 | 0.01× | 94.1655 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 3.1088 122.3452 0.7299 | · | · | · |
Reduction¶
VarMeanFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-var-mean - W3x: [2048, 4096]dtype=f16
hidden-state-var-mean - W4x: [64, 32768]dtype=bf16
long-seq-var-mean
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 4.45× | 0.0264 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1175 0.1090 | · | · | · |
| W2 | 2.69× | 0.0535 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1442 0.1361 | · | · | · |
| W3 | 2.80× | 0.0512 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.1427 0.1368 | · | · | · |
| W4 | 2.92× | 0.0275 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis | 0.0810 0.0746 | · | · | · |
Normalization¶
AdaLayerNormZeroFwd¶
- W1x: [1024, 1152]dtype=bf16
dit-xl-2 - W2x: [1024, 1152]dtype=f16
dit-xl-2 - W3x: [1, 4096]dtype=bf16
llama-8b-decode - W4x: [2048, 4096]dtype=bf16
llama-8b-prefill - W5x: [2048, 4096]dtype=f16
llama-8b-prefill
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 2.33× | 0.0269 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0622 0.0549 | · | · | · |
| W2 | 2.38× | 0.0272 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.0617 0.0578 | · | · | · |
| W3 | 3.12× | 0.004 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basistilelang-mlir-ascend:high_perf tilelang-mlir-ascend ⓘ eager · n=5 ⓘ | 0.0123 0.0110 0.00424 | · | · | · |
| W4 | 1.29× | 0.1231 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1593 0.1464 | · | · | · |
| W5 | 1.51× | 0.1205 | torch_compile_aclgraph torch_compiletorch_npu eager vendor library basis ⓘ | 0.1816 0.1737 | · | · | · |
InfNormFwd¶
- W1x: [4, 128, 4096]dtype=f16dim=[0,2]
3d-multidim-reduce - W2x: [2048, 4096]dtype=bf16
hidden-state-inf - W3x: [2048, 4096]dtype=f16
hidden-state-inf - W4x: [64, 32768]dtype=bf16
long-seq-inf
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.87× | 0.0357 | Torch-NPU eager reference | 0.0310 | · | · | · |
| W2 | 1.64× | 0.0360 | Torch-NPU eager reference | 0.0591 | · | · | · |
| W3 | 1.57× | 0.0375 | Torch-NPU eager reference | 0.0589 | · | · | · |
| W4 | 1.55× | 0.0187 | Torch-NPU eager reference | 0.0290 | · | · | · |
Pooling¶
AdaptiveMaxPool2dIndicesFwd¶
input: [N, C, H_in, W_in]
- W1dtype=f16N=8C=2048H_in=7W_in=7output_size=[1,1]
global-1x1 - W2dtype=bf16N=2C=64H_in=55W_in=57output_size=[7,7]
nondiv-7x7 - W3dtype=f16N=2C=128H_in=56W_in=56output_size=[6,6]
spp-6x6
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.06× | 1.7141 | Torch-NPU eager reference | 0.0973 | · | · | · |
| W2 | 0.10× | 1.3225 | Torch-NPU eager reference | 0.1296 | · | · | · |
| W3 | 0.05× | 3.0819 | Torch-NPU eager reference | 0.1675 | · | · | · |
MaxPool3dIndicesFwd¶
input: [N, C, D_in, H_in, W_in]
- W1dtype=f16N=8C=64D_in=16H_in=112W_in=112kernel_size=[1,2,2]stride=[1,2,2]padding=[0,0,0]
c3d-pool1 - W2dtype=f16N=4C=128D_in=16H_in=56W_in=56kernel_size=[2,2,2]stride=[2,2,2]padding=[0,0,0]
c3d-pool2 - W3dtype=bf16N=2C=64D_in=32H_in=112W_in=112kernel_size=[3,3,3]stride=[2,2,2]padding=[1,1,1]ceil_mode=true
medicalnet-stem
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.00× | 240.7824 | torch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.9711 0.7434 | · | · | · |
| W2 | 0.00× | 64.8989 | torch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 0.3008 0.1850 | · | · | · |
| W3 | 0.00× | 558.8976 | torch_compile_ge torch_compiletorch_npu eager vendor library basis ⓘ | 2.4289 1.2630 | · | · | · |
Positional Encoding¶
RopeNeoxPositionIdsFwd¶
x: [num_tokens, num_heads, head_dim]position_ids: [num_tokens]
num_heads=32head_dim=128
- W1dtype=f16num_tokens=2048max_position=4096
position-ids-s2k-h32-d128 - W2dtype=bf16num_tokens=4096max_position=8192
position-ids-s4k-h32-d128
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.85× | 0.1086 | ops-transformer-b2:aclnnRopeWithSinCosCache open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.2000 0.2979 0.2835 | · | · | · |
| W2 | 1.91× | 0.2024 | ops-transformer-b2:aclnnRopeWithSinCosCache open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.3875 0.4447 0.4719 | · | · | · |
RopeNeoxFwd¶
- W1dtype=f16seq_len=2048head_dim=64
neox-1d-2k-d64 - W2dtype=bf16seq_len=4096head_dim=128
neox-1d-4k-d128 - W3dtype=f16batch=2seq_len=2048num_heads=32head_dim=128layout=2d
neox-2d-b2-s2k-h32-d128
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.14× | 0.0143 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.0118 0.0698 0.0638 | · | · | · |
| W2 | 0.96× | 0.0173 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.0105 0.0757 0.0719 | · | · | · |
| W3 | 0.91× | 0.1145 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.0949 0.3815 0.3954 | · | · | · |
RopeLlama31Fwd¶
seq_len=8192head_dim=128
- W1dtype=bf16
llama31-1d-8k-d128 - W2dtype=f16batch=1num_heads=32layout=2d
llama31-2d-b1-s8k-h32-d128
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.04× | 0.0206 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.0155 0.0921 0.0780 | · | · | · |
| W2 | 0.86× | 0.2246 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library ⓘ | 0.1874 1.0653 1.1926 | · | · | · |
RopeYarnFwd¶
seq_len=8192head_dim=128
- W1dtype=bf16
yarn-1d-8k-d128 - W2dtype=f16batch=1num_heads=32layout=2d
yarn-2d-b1-s8k-h32-d128
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.99× | 0.0209 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.0152 0.0925 0.0772 | · | · | · |
| W2 | 0.86× | 0.2286 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1898 1.0504 1.1845 | · | · | · |
RopeLongRopeFwd¶
seq_len=8192head_dim=128
- W1dtype=bf16
longrope-1d-8k-d128 - W2dtype=f16batch=1num_heads=32layout=2d
longrope-2d-b1-s8k-h32-d128
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.94× | 0.0203 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.0150 0.0920 0.0790 | · | · | · |
| W2 | 0.86× | 0.2300 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1908 1.0408 1.2095 | · | · | · |
RopeNonNeoxFwd¶
seq_len=2048
- W1dtype=f16head_dim=64
non-neox-1d-2k-d64 - W2dtype=bf16batch=2num_heads=32head_dim=128layout=2d
non-neox-2d-b2-s2k-h32-d128
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 0.91× | 0.0170 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.00813 0.0938 0.1526 | · | · | · |
| W2 | 0.72× | 0.1965 | ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basistorch_compile_aclgraph torch_compiletorch_npu eager vendor library | 0.1358 0.4402 0.5316 | · | · | · |
MHC¶
MHCPostFwd¶
x_layer_out: [batch, c_x]h_post: [batch, n_expand], f32x_res: [batch, 10240]
dtype=bf16batch=4c_x=2560n_expand=4
- W1
post-large
x_layer_out: [batch, c_x]h_post: [batch, n_expand], f32x_res: [batch, 7680]
dtype=bf16batch=2c_x=1920n_expand=4
- W2
post-medium
x_layer_out: [batch, c_x]h_post: [batch, n_expand], f32x_res: [batch, 5120]
dtype=bf16batch=1c_x=1280n_expand=4
- W3
post-small
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name · tier | ms | TFLOP/s | of ceiling | by | |
| W1 | 3.22× | 0.0112 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0362 0.0278 0.0296 | · | · | · |
| W2 | 2.82× | 0.0085 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0240 0.0195 0.0170 | · | · | · |
| W3 | 3.83× | 0.0045 | torch_compile_aclgraph torch_compiletorch_compile_ge torch_compiletorch_npu eager vendor library basis | 0.0195 0.0123 0.0123 | · | · | · |
MHCPreFwd¶
phi: [10240, 24], f32x: [batch, 10240]b: [24], f32
dtype=bf16batch=4alpha_pre=1alpha_post=1alpha_res=1sinkhorn_repeat=1
- W1
pre-large
phi: [7680, 24], f32x: [batch, 7680]b: [24], f32
dtype=bf16batch=2alpha_pre=1alpha_post=1alpha_res=1sinkhorn_repeat=1
- W2
pre-medium
phi: [5120, 24], f32x: [batch, 5120]b: [24], f32
dtype=bf16batch=1alpha_pre=1alpha_post=1alpha_res=1sinkhorn_repeat=1
- W3
pre-small
| Workload | Ratio | Device time | Alternatives | Throughput | SOL | Bound | |
|---|---|---|---|---|---|---|---|
| alt / ours | ms | name | ms | TFLOP/s | of ceiling | by | |
| W1 | 1.06× | 0.2026 | Torch-NPU eager exact reference | 0.2150 | · | · | · |
| W2 | 1.03× | 0.1253 | Torch-NPU eager exact reference | 0.1295 | · | · | · |
| W3 | 1.14× | 0.0757 | Torch-NPU eager exact reference | 0.0860 | · | · | · |