Skip to content

Second-phase operators

50 ops, 236 workloads — Attention 9 · Linear Attention / SSM 2 · Mixture of Experts 1 · Quantization 3 · Elementwise 22 · Reduction 1 · Normalization 2 · Pooling 2 · Positional Encoding 6 · MHC 2.

One table per op, one row per workload. Ratio is the baseline device time divided by ours in the same measurement regime, so green is faster than it, plain is level with it, red is slower. Times are in ms. How these numbers are taken.

A separate set, and a separate denominator

These operators are not in oplist150, the first-phase acceptance list (docs/reports/R267-data/oplist150.json) that the overview and the per-family data pages are scoped to and that progress is reported against. They are shown here because the harness already measured a real opponent against them, not because the first phase grew.

Nothing on this page is added to any count on the other pages. The overview's coverage figures, the per-family tallies and the attainment counts are all first-phase only, and stay first-phase only. This page's own tally, below, counts these operators and nothing else.

An operator reaches this page only when the snapshot publishes a ratio for it. One with no timed opponent is left out rather than shown as a row of blanks.

Attention

GroupedQueryAttentionDecodePagedWithKVCacheFwd

dtype=f16

  • W1q: [4, 128, 128]kv: [4096, 8, 128]page_size=256serving-405b-p256
  • W2q: [8, 64, 128]kv: [4096, 8, 128]page_size=256serving-70b-p256
  • W3q: [8, 64, 128]kv: [4096, 8, 128]page_size=64serving-70b-p64
  • W4q: [8, 32, 128]kv: [32768, 8, 128]page_size=64serving-8b-long-p64
  • W5q: [32, 32, 128]kv: [4096, 8, 128]page_size=256serving-8b-p256
  • W6q: [32, 32, 128]kv: [4096, 8, 128]page_size=64serving-8b-p64
  • W7q: [64, 32, 128]kv: [2048, 8, 128]page_size=64throughput-8b-p64
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.45×33.4937torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager python reference vendor library
48.4476
48.5021
48.5410
···
W21.19×33.4759torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager python reference vendor library
39.9049
39.9079
44.9059
···
W3—·torch_compile_ge torch_compile basis
torch_npu eager python reference vendor library
39.6264
44.8341
···
W4—·torch_compile_ge torch_compile basis
torch_npu eager python reference vendor library
136.3076
165.6499
···
W51.18×65.3830torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager python reference vendor library
77.1264
77.1175
87.2837
···
W6—·torch_compile_ge torch_compile basis
torch_npu eager python reference vendor library
76.0240
88.2759
···
W7—·torch_compile_ge torch_compile basis
torch_npu eager python reference vendor library
144.2580
161.4332
···

GroupedQueryAttentionSlidingWindowVarlenFwd

dtype=f16batch=4dim=128

  • W1total_q=8192total_k=8192heads=64/8window_size_left=1024max_seqlen_q=2048llama-70b-long-w1024
  • W2total_q=2048total_k=2048heads=64/8window_size_left=256max_seqlen_q=512llama-70b-short-w256
  • W3total_q=8192total_k=8192heads=32/8window_size_left=1024max_seqlen_q=2048llama-8b-long-w1024
  • W4total_q=2048total_k=2048heads=32/8window_size_left=256max_seqlen_q=512llama-8b-short-w256
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.81×6.2363torch_compile_aclgraph torch_compile
torch_npu eager python reference vendor library basis
5.0464
5.0319
···
W21.16×0.5373torch_compile_aclgraph torch_compile basis
torch_npu eager python reference vendor library
0.6249
0.6279
···
W30.80×3.3505torch_compile_aclgraph torch_compile basis
torch_npu eager python reference vendor library
2.6565
2.7251
···
W41.26×0.3385torch_compile_aclgraph torch_compile
torch_npu eager python reference vendor library basis
0.4236
0.4211
···

MultiHeadAttentionFwd

  • W1q, kv: [1, 2048, 64, 128]dtype=bf16llama-70b-long
  • W2q, kv: [1, 2048, 64, 128]dtype=f16llama-70b-long
  • W3q, kv: [2, 512, 64, 128]dtype=bf16llama-70b-short
  • W4q, kv: [2, 512, 64, 128]dtype=f16llama-70b-short
  • W5q, kv: [2, 2048, 32, 128]dtype=bf16llama-8b-long
  • W6q, kv: [2, 2048, 32, 128]dtype=f16llama-8b-long
  • W7q, kv: [4, 512, 32, 128]dtype=bf16llama-8b-short
  • W8q, kv: [4, 512, 32, 128]dtype=f16llama-8b-short
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.42×1.5755ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager SDPA vendor library
0.6620
0.8296
0.8330
0.8325
···
W20.42×1.5684ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.6521
0.8214
0.8235
···
W30.63×0.2579ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.1611
0.2915
0.2884
···
W40.65×0.2534ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.1608
0.2904
0.2855
···
W50.42×1.5901ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.6626
0.8350
0.8306
···
W60.41×1.5860ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.6498
0.8173
0.8165
···
W70.64×0.2560ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.1616
0.2949
0.2945
···
W80.66×0.2555ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.1638
0.2984
0.2961
···

GroupedQueryAttentionFwd

dtype=f16

  • W1q: [1, 2048, 64, 128]kv: [1, 2048, 8, 128]llama-70b-long
  • W2q: [2, 512, 64, 128]kv: [2, 512, 8, 128]llama-70b-short
  • W3q: [2, 2048, 32, 128]kv: [2, 2048, 8, 128]llama-8b-long
  • W4q: [4, 512, 32, 128]kv: [4, 512, 8, 128]llama-8b-short
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.41×1.5300ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.6172
0.7180
0.7185
···
W20.60×0.2390ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.1411
0.2612
0.2616
···
W30.41×1.5433ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.6211
0.7411
0.7349
···
W40.60×0.2439ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library
0.1449
0.2695
0.2679
···

GroupedQueryAttentionSlidingWindowFwd

dtype=f16

  • W1q: [1, 2048, 64, 128]kv: [1, 2048, 8, 128]window_size_left=1024llama-70b-long-w1024
  • W2q: [2, 512, 64, 128]kv: [2, 512, 8, 128]window_size_left=256llama-70b-short-w256
  • W3q: [2, 2048, 32, 128]kv: [2, 2048, 8, 128]window_size_left=1024llama-8b-long-w1024
  • W4q: [4, 512, 32, 128]kv: [4, 512, 8, 128]window_size_left=256llama-8b-short-w256
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.30×1.8444ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager python reference vendor library
0.5539
1.5036
1.5085
···
W20.43×0.3362ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager python reference vendor library
0.1391
0.3902
0.3925
···
W30.30×1.8560ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager python reference vendor library
0.5573
1.5281
1.5203
···
W40.43×0.3429ops-transformer-b1:aclnnFlashAttentionScore open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager python reference vendor library
0.1434
0.3946
0.3921
···

MultiHeadAttentionDecodeWithKVCacheFwd

  • W1q: [4, 1, 64, 128]kv: [4, 32768, 64, 128]dtype=bf16llama-70b-32k
  • W2q: [4, 1, 64, 128]kv: [4, 32768, 64, 128]dtype=f16llama-70b-32k
  • W3q: [16, 1, 64, 128]kv: [16, 4096, 64, 128]dtype=bf16llama-70b-4k
  • W4q: [16, 1, 64, 128]kv: [16, 4096, 64, 128]dtype=f16llama-70b-4k
  • W5q: [8, 1, 32, 128]kv: [8, 32768, 32, 128]dtype=bf16llama-8b-32k
  • W6q: [8, 1, 32, 128]kv: [8, 32768, 32, 128]dtype=f16llama-8b-32k
  • W7q: [32, 1, 32, 128]kv: [32, 4096, 32, 128]dtype=bf16llama-8b-4k
  • W8q: [32, 1, 32, 128]kv: [32, 4096, 32, 128]dtype=f16llama-8b-4k
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.13×139.8074torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager SDPA vendor library basis
18.6446
19.0927
18.4155
···
W20.13×139.8650torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library basis
18.7369
18.6160
···
W30.14×65.0734torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library basis
9.2641
9.2630
···
W40.14×65.1240torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library basis
9.2516
9.2351
···
W50.15×139.7050torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA vendor library
20.7512
20.7583
···
W60.15×139.7111torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA vendor library
20.6946
20.7183
···
W70.14×65.0439torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library basis
9.0760
9.0411
···
W80.14×65.1153torch_compile_aclgraph torch_compile
torch_npu eager SDPA vendor library basis
9.0676
9.0475
···

GroupedQueryAttentionBwd

dtype=f16

  • W1q: [1, 2048, 64, 128]kv: [1, 2048, 8, 128]llama-70b-long
  • W2q: [2, 512, 64, 128]kv: [2, 512, 8, 128]llama-70b-short
  • W3q: [2, 2048, 32, 128]kv: [2, 2048, 8, 128]llama-8b-long
  • W4q: [4, 512, 32, 128]kv: [4, 512, 8, 128]llama-8b-short
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.01×270.8478torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA autograd vendor library
2.3700
2.3945
···
W20.02×44.8023torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA autograd vendor library
0.7984
0.8086
···
W30.01×272.6906torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA autograd vendor library
2.4531
2.4635
···
W40.02×45.1514torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA autograd vendor library
0.7960
0.8096
···

MultiHeadAttentionBwd

dtype=f16

  • W1q, kv: [1, 2048, 64, 128]llama-70b-long
  • W2q, kv: [2, 512, 64, 128]llama-70b-short
  • W3q, kv: [2, 2048, 32, 128]llama-8b-long
  • W4q, kv: [4, 512, 32, 128]llama-8b-short
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.01×365.0621torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA autograd vendor library
2.9143
2.9274
···
W20.01×91.6334torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA autograd vendor library
0.7804
0.7946
···
W30.01×365.5979torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA autograd vendor library
2.8984
2.9120
···
W40.01×91.8776torch_compile_aclgraph torch_compile basis
torch_npu eager SDPA autograd vendor library
0.7937
0.7959
···

MultiHeadLatentAttentionDecodeWithKVCacheFwd

dtype=f16pe_dim=64

  • W1q: [8, 128, 128]q_pe: [8, 128, 64]kv: [8, 32768, 1, 128]k_pe: [8, 32768, 1, 64]deepseek-v2-32k
  • W2q: [32, 128, 128]q_pe: [32, 128, 64]kv: [32, 4096, 1, 128]k_pe: [32, 4096, 1, 64]deepseek-v2-4k
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.00×521.9301torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager python reference vendor library basis
0.6825
0.6180
0.6683
···
W20.00×260.9430torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager python reference vendor library basis
0.3005
0.2662
0.2871
···

Linear Attention / SSM

EngramGateConvBwd

dY, H, k, v, vhat: [M, seq_len, d]rms_w_h, rms_w_v: [d]conv_w: [4, d]alpha, rrms_h, rrms_k, rrms_v: [M, seq_len], f32

  • W1dtype=bf16M=1seq_len=128d=256bwd-b1-s128-d256
  • W2dtype=f16M=1seq_len=32d=256bwd-b1-s32-d256
  • W3dtype=f16M=2seq_len=64d=512bwd-b2-s64-d512
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W111.01×0.0831Torch-NPU eager exact reference0.9155···
W217.36×0.0408Torch-NPU eager exact reference0.7074···
W312.66×0.0825Torch-NPU eager exact reference1.0442···

EngramGateConvFwd

H, k, v: [M, seq_len, d]rms_w_h, rms_w_v: [d]conv_w: [4, d]

  • W1dtype=bf16M=1seq_len=128d=256fwd-b1-s128-d256
  • W2dtype=f16M=1seq_len=32d=256fwd-b1-s32-d256
  • W3dtype=f16M=2seq_len=64d=512fwd-b2-s64-d512
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W18.71×0.0410Torch-NPU eager exact reference0.3571···
W29.95×0.0280Torch-NPU eager exact reference0.2785···
W38.76×0.0452Torch-NPU eager exact reference0.3962···

Mixture of Experts

MoeUnpermuteFwd

dtype=bf16top_k=8

  • W1total_tokens=1hidden_size=7168large-hidden-decode
  • W2total_tokens=512hidden_size=7168large-hidden-medium
  • W3total_tokens=4096hidden_size=7168large-hidden-prefill
  • W4total_tokens=32hidden_size=7168large-hidden-small
  • W5total_tokens=1hidden_size=3072small-hidden-decode
  • W6total_tokens=512hidden_size=3072small-hidden-medium
  • W7total_tokens=4096hidden_size=3072small-hidden-prefill
  • W8total_tokens=32hidden_size=3072small-hidden-small
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.84×0.00625ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.00425
0.0445
0.0245
···
W20.51×0.1747ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.0870
0.3869
0.3757
···
W30.41×1.3121ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.5310
2.8803
2.8481
···
W40.64×0.0180ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.00975
0.0600
0.0454
···
W50.90×0.005ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.00313
0.0380
0.0180
···
W60.35×0.1300ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.0420
0.1510
0.1355
···
W70.29×0.9976ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.2871
1.2434
1.2379
···
W80.64×0.0152ops-transformer-b2:aclnnMoeTokenUnpermute open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.00675
0.0561
0.0400
···

Quantization

FP8LightningIndexerFwd

index_q: [batch, seq_len, heads, index_dim]index_k: [batch, seq_len_kv, kv_group, index_dim]weights: [seq_len, heads], f32cu_seqlen_ks, cu_seqlen_ke: [seq_len], i32index_k_scale: [batch, seq_len_kv, kv_group], f32

batch=1seq_len=8192heads=32index_dim=64seq_len_kv=32768kv_group=1

  • W1dtype=bf16lightning-indexer-s8k-h32-d64
  • W2dtype=fp8e4m3lightning-indexer-s8k-h32-d64

index_k_scale: [batch, seq_len_kv, kv_group], f32

dtype=fp8e4m3batch=1seq_len_kv=32768kv_group=1seq_len=8192heads=32index_dim=64

  • W3lightning-indexer-s8k-h32-d64-kscale
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W15.88×24.0535torch_npu eager python reference vendor library basis ⓘ141.4828···
W25.82×24.2429torch_npu eager python reference vendor library basis ⓘ141.1920···
W35.70×25.1346torch_npu eager python reference vendor library basis ⓘ143.3777···

BmmFp8NKFwd

dtype=fp8e4m3

  • W1b=32m=128n=128k=2048mha-decode-b32-pv-per-tensor
  • W2b=64m=128n=2048k=128mha-decode-b64-qk-per-tensor
  • W3b=128m=512n=512k=2048moe-prefill-b128-per-tensor
  • W4b=4m=1024n=1024k=1024square-b4-1k-per-tensor
  • W5b=8m=2048n=2048k=2048square-b8-2k-per-tensor
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W14.83×0.1133torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.5479
0.5954
···
W23.65×0.2238torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.8139
0.9902
···
W39.28×2.3949torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
22.3016
22.2772
···
W42.42×0.1383torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.3374
0.3259
···
W53.48×1.4383torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
5.0110
5.5925
···

BmmFp8KNFwd

dtype=fp8e4m3

  • W1b=32m=128n=128k=2048mha-decode-b32-pv-per-tensor
  • W2b=64m=128n=2048k=128mha-decode-b64-qk-per-tensor
  • W3b=128m=512n=512k=2048moe-prefill-b128-per-tensor
  • W4b=4m=1024n=1024k=1024square-b4-1k-per-tensor
  • W5b=8m=2048n=2048k=2048square-b8-2k-per-tensor
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W14.12×0.1305torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.5369
0.5880
···
W23.87×0.2070torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
0.8036
0.9409
···
W39.28×2.3847torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
22.2211
22.1105
···
W42.36×0.1400torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.3287
0.3186
···
W53.50×1.4418torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
5.0496
5.6059
···

Elementwise

ClampScalarFwd

min=-0.5max=0.5

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W14.18×0.0691ClipByValueV2_e87c9546cf2804355cf0d9445a2db01a_high_performance_2200100000.2890···
W22.23×0.0683ClipByValueV2_1056052785b43db3ca4234fbf5e08c0e_high_performance_2200100000.1522···
W32.50×0.1150ClipByValueV2_54f1d1caf8fe36c496f699fefe98fbca_high_performance_2200100000.2878···
W44.73×0.9459ClipByValueV2_e87c9546cf2804355cf0d9445a2db01a_high_performance_2200100004.4703···
W52.35×0.9565ClipByValueV2_1056052785b43db3ca4234fbf5e08c0e_high_performance_2200100002.2517···

BitwiseNotFwd

  • W1input: [4096, 4096]dtype=i32elementwise-16M
  • W2input: [4096, 4096]dtype=i64elementwise-16M
  • W3input: [16384, 16384]dtype=i32elementwise-256M
  • W4input: [4097]dtype=i32legacy-probe-1
  • W5input: [4097]dtype=i64legacy-probe-2
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W14.53×5.0855torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
22.9669
22.9959
···
W26.08×5.1049torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
31.4594
31.4969
···
W34.53×80.5434torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library ⓘ
364.8630
364.9903
···
W40.77×0.0410torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0325
0.0284
···
W50.83×0.0431torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0333
0.0297
···

AlibiFwd

num_heads=32

  • W1dtype=bf16seq_len=2048llama-prefill-2k
  • W2dtype=f16seq_len=2048llama-prefill-2k
  • W3dtype=bf16seq_len=4096llama-prefill-4k
  • W4dtype=f16seq_len=4096llama-prefill-4k
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W12.25×0.5493Torch-NPU eager reference1.2385···
W22.25×0.5505Torch-NPU eager reference1.2361···
W32.16×2.1581Torch-NPU eager reference4.6643···
W42.16×2.1609Torch-NPU eager reference4.6753···

PreluFwd

  • W1input: [1, 4, 1025]weight: [4]dtype=f16channel-tail-nondivisible
  • W2input: [16, 256, 56, 56]weight: [256]dtype=bf16cnn-feat-per-channel
  • W3input: [16, 512, 28, 28]weight: [512]dtype=bf16cnn-feat-per-channel-deep
  • W4input: [16, 512, 28, 28]weight: [512]dtype=f16cnn-feat-per-channel-deep
  • W5input: [16, 256, 56, 56]weight: [256]dtype=f16cnn-feat-per-channel
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W10.61×0.0248Torch-NPU eager exact reference0.0150···
W22.43×0.0620Torch-NPU eager exact reference0.1509···
W32.41×0.0379Torch-NPU eager exact reference0.0912···
W42.41×0.0380Torch-NPU eager exact reference0.0917···
W52.30×0.0670Torch-NPU eager exact reference0.1541···

MaskedFillScalarFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W11.67×0.0843Torch-NPU eager exact reference0.1410···
W21.25×0.0848Torch-NPU eager exact reference0.1060···
W31.38×0.1395Torch-NPU eager exact reference0.1925···
W42.18×1.1220Torch-NPU eager exact reference2.4454···
W51.85×1.1454Torch-NPU eager exact reference2.1162···

LogicalOrFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=boolcnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W4input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W5input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=boolhidden-state-prefill
  • W7input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W8input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.73×0.0612torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0998
0.0998
0.0998
···
W20.94×0.0452torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library
0.0359
0.0356
0.0360
···
W31.83×0.0512torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0835
0.0833
0.0835
···
W41.60×0.0895torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1360
0.1363
0.1360
···
W51.71×0.0576torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0990
0.0931
0.0830
···
W61.00×0.0356torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0355
0.0325
0.0297
···
W71.53×0.0559torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0828
0.0770
0.0674
···
W81.37×0.1034torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1415
0.1325
0.1224
···

NeFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.81×0.0658torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.1077
0.1077
0.1077
···
W21.01×0.0563torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0560
0.0548
0.0530
···
W31.29×0.0945torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.1163
0.1160
0.1163
···
W41.37×0.0644torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0885
0.0794
0.0793
···
W50.96×0.0616torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0592
0.0607
0.0553
···
W61.16×0.1062torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.1217
0.1123
0.1153
···

GtFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.22×0.0595torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0653
0.0650
0.0651
···
W21.06×0.0512torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0540
0.0678
0.0464
···
W31.10×0.0835torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0833
0.0833
0.0831
···
W41.15×0.0545torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0626
0.0550
0.0548
···
W50.93×0.0568torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0530
0.0510
0.0485
···
W61.04×0.0943torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0980
0.0998
0.0935
···

LtFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.21×0.0578torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0602
0.0602
0.0600
···
W20.98×0.0455torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0480
0.0633
0.0415
···
W31.18×0.0789torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0789
0.0788
0.0786
···
W41.16×0.0504torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0580
0.0532
0.0493
···
W50.95×0.0478torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0451
0.0430
0.0393
···
W61.01×0.0917torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0960
0.0944
0.0902
···

LeFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.16×0.0553torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.0573
0.0575
0.0574
···
W21.01×0.0457torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0460
0.0558
0.0405
···
W31.17×0.0756torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0747
0.0751
0.0746
···
W41.15×0.0483torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0553
0.0503
0.0480
···
W50.96×0.0483torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0465
0.0455
0.0428
···
W61.03×0.0880torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0877
0.0876
0.0846
···

GeFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.16×0.0633torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.0673
0.0673
0.0674
···
W21.10×0.0452torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0498
0.0658
0.0430
···
W31.08×0.0830torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0833
0.0835
0.0833
···
W41.10×0.0537torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0575
0.0537
0.0514
···
W50.96×0.0505torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0498
0.0491
0.0447
···
W61.04×0.0958torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.1000
0.1010
0.0970
···

EqFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.13×0.0665torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0653
0.0653
0.0651
···
W21.03×0.0560torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0573
0.0575
0.0527
···
W31.09×0.0932torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0927
0.0927
0.0925
···
W41.10×0.0633torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0700
0.0633
0.0636
···
W50.94×0.0620torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0587
0.0580
0.0550
···
W61.04×0.1057torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.1082
0.1065
0.1045
···

NanToNumFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W11.11×0.0798NanToNum_71a4100bd8f51c70f591051c16121929_high_performance_2100100000.0885···
W20.86×0.0786NanToNum_ccb78600df288a689f3cc4abbf7e007e_high_performance_2100100000.0673···
W30.95×0.1295NanToNum_5703d6e339578b11bc6d39a70205aa62_high_performance_2100100000.1229···
W41.08×1.0601NanToNum_71a4100bd8f51c70f591051c16121929_high_performance_2100100001.1493···
W50.88×1.0551NanToNum_ccb78600df288a689f3cc4abbf7e007e_high_performance_2100100000.9250···

SinusoidalFwd

d_model=4096

  • W1dtype=bf16seq_len=2048transformer-2k-4k
  • W2dtype=f16seq_len=2048transformer-2k-4k
  • W3dtype=bf16seq_len=4096transformer-4k-4k
  • W4dtype=f16seq_len=4096transformer-4k-4k
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W11.02×0.4160Torch-NPU eager reference0.4228···
W21.03×0.4193Torch-NPU eager reference0.4328···
W30.91×0.8217Torch-NPU eager reference0.7468···
W40.91×0.8220Torch-NPU eager reference0.7458···

SeluFwd

  • W1input: [2048, 4096]dtype=bf16snn-fc
  • W2input: [2048, 4096]dtype=f16snn-fc
  • W3input: [2048, 8192]dtype=bf16snn-fc-wide
  • W4input: [2048, 8192]dtype=f16snn-fc-wide
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.84×0.0440torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0367
0.0367
0.0318
···
W20.78×0.0460torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0360
0.0376
0.0323
0.0296
···
W30.84×0.0779torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.0663
0.0661
0.0612
···
W40.84×0.0770torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:npuir_active tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0643
0.0648
0.0607
0.0586
···

FloorDivideFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=bf16cnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f16cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=f32cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=bf16hidden-state-prefill
  • W5input, other: [2048, 4096]dtype=f16hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=f32hidden-state-prefill
  • W7input, other: [4097]dtype=bf16nondiv-tail
  • W8input, other: [4097]dtype=f16nondiv-tail
  • W9input, other: [4097]dtype=f32nondiv-tail
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.52×0.1578torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0814
0.0650
···
W20.44×0.1555torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0688
0.0586
···
W30.61×0.1580torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0993
0.0917
···
W40.61×0.1052torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0620
0.0563
···
W50.60×0.1054torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0639
0.0573
···
W60.87×0.1227torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1062
0.1015
···
W70.69×0.00325torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.00225
0.0015
···
W80.75×0.003torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.00225
0.0015
···
W90.75×0.003torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0015
0.0015
···

BitwiseXorFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=boolcnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i32cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i64cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=boolhidden-state-prefill
  • W5input, other: [2048, 4096]dtype=i32hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=i64hidden-state-prefill
  • W7input, other: [4097]dtype=i32tail-nondivisible
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.01×4.3217torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0602
0.0455
···
W20.02×7.0530torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1075
0.0986
···
W30.02×10.1686torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1926
0.1872
···
W41.15×0.0333torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0379
0.0306
···
W51.04×0.1020torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1062
0.1027
···
W61.07×0.1956torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.2102
0.2025
···
W70.89×0.00225torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.002
0.0015
···

BitwiseAndFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=boolcnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i32cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i64cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=boolhidden-state-prefill
  • W5input, other: [2048, 4096]dtype=i32hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=i64hidden-state-prefill
  • W7input, other: [4097]dtype=i32tail-nondivisible
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.01×4.3252torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0483
0.0350
···
W20.01×7.0652torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0993
0.0912
···
W30.02×10.1616torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1905
0.1815
···
W41.08×0.0328torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0362
0.0312
···
W51.00×0.1051torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1057
0.1025
···
W61.04×0.1993torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.2055
0.2004
···
W71.06×0.002torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.00225
0.0015
···

BitwiseOrFwd

  • W1input: [16, 256, 56, 56]other: [256, 1, 1]dtype=boolcnn-feat-broadcast
  • W2input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i32cnn-feat-broadcast
  • W3input: [16, 256, 56, 56]other: [256, 1, 1]dtype=i64cnn-feat-broadcast
  • W4input, other: [2048, 4096]dtype=boolhidden-state-prefill
  • W5input, other: [2048, 4096]dtype=i32hidden-state-prefill
  • W6input, other: [2048, 4096]dtype=i64hidden-state-prefill
  • W7input, other: [4097]dtype=i32tail-nondivisible
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.01×4.3249torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0504
0.0357
···
W20.01×7.0569torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0998
0.0920
···
W30.02×10.1641torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1883
0.1815
···
W41.07×0.0318torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0339
0.0296
···
W51.01×0.1040torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1070
0.1010
···
W61.04×0.1968torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.2065
0.1980
···
W71.00×0.002torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.00225
0.0015
···

IsinfFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.02×5.9417torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0930
0.0571
0.0809
···
W20.01×6.0650torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0770
0.0555
0.0678
···
W30.02×5.9460torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.1360
0.1035
0.1303
···
W40.02×94.2352torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
1.7075
1.7066
1.7060
···
W50.02×94.1644torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
1.6007
0.7289
1.5947
···

IsnanFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.02×6.5216torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1148
0.1067
···
W20.01×8.8604torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0587
0.0542
···
W30.02×6.5269torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1309
0.1212
···
W40.02×103.4591torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
1.6560
1.6482
···
W50.01×103.3778torch_compile_aclgraph torch_compile basis
torch_npu eager vendor library
0.7666
0.7720
···

IsfiniteFwd

  • W1input: [4096, 4096]dtype=bf16elementwise-16M
  • W2input: [4096, 4096]dtype=f16elementwise-16M
  • W3input: [4096, 4096]dtype=f32elementwise-16M
  • W4input: [16384, 16384]dtype=bf16elementwise-256M
  • W5input: [16384, 16384]dtype=f16elementwise-256M
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.01×5.9420torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.2279
7.7950
0.0555
···
W20.01×6.0667torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.1160
7.7988
0.0545
···
W30.02×5.9465torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.2697
7.8072
0.0985
···
W40.01×94.2362torch_compile_aclgraph torch_compile basis
torch_compile_ge torch_compile
torch_npu eager vendor library ⓘ
0.7360
0.7370
0.7366
···
W50.01×94.1655torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
3.1088
122.3452
0.7299
···

Reduction

VarMeanFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-var-mean
  • W3x: [2048, 4096]dtype=f16hidden-state-var-mean
  • W4x: [64, 32768]dtype=bf16long-seq-var-mean
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W14.45×0.0264torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1175
0.1090
···
W22.69×0.0535torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1442
0.1361
···
W32.80×0.0512torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.1427
0.1368
···
W42.92×0.0275torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
0.0810
0.0746
···

Normalization

AdaLayerNormZeroFwd

  • W1x: [1024, 1152]dtype=bf16dit-xl-2
  • W2x: [1024, 1152]dtype=f16dit-xl-2
  • W3x: [1, 4096]dtype=bf16llama-8b-decode
  • W4x: [2048, 4096]dtype=bf16llama-8b-prefill
  • W5x: [2048, 4096]dtype=f16llama-8b-prefill
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W12.33×0.0269torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0622
0.0549
···
W22.38×0.0272torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.0617
0.0578
···
W33.12×0.004torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis
tilelang-mlir-ascend:high_perf tilelang-mlir-ascend ⓘ eager · n=5 ⓘ
0.0123
0.0110
0.00424
···
W41.29×0.1231torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1593
0.1464
···
W51.51×0.1205torch_compile_aclgraph torch_compile
torch_npu eager vendor library basis ⓘ
0.1816
0.1737
···

InfNormFwd

  • W1x: [4, 128, 4096]dtype=f16dim=[0,2]3d-multidim-reduce
  • W2x: [2048, 4096]dtype=bf16hidden-state-inf
  • W3x: [2048, 4096]dtype=f16hidden-state-inf
  • W4x: [64, 32768]dtype=bf16long-seq-inf
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W10.87×0.0357Torch-NPU eager reference0.0310···
W21.64×0.0360Torch-NPU eager reference0.0591···
W31.57×0.0375Torch-NPU eager reference0.0589···
W41.55×0.0187Torch-NPU eager reference0.0290···

Pooling

AdaptiveMaxPool2dIndicesFwd

input: [N, C, H_in, W_in]

  • W1dtype=f16N=8C=2048H_in=7W_in=7output_size=[1,1]global-1x1
  • W2dtype=bf16N=2C=64H_in=55W_in=57output_size=[7,7]nondiv-7x7
  • W3dtype=f16N=2C=128H_in=56W_in=56output_size=[6,6]spp-6x6
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W10.06×1.7141Torch-NPU eager reference0.0973···
W20.10×1.3225Torch-NPU eager reference0.1296···
W30.05×3.0819Torch-NPU eager reference0.1675···

MaxPool3dIndicesFwd

input: [N, C, D_in, H_in, W_in]

  • W1dtype=f16N=8C=64D_in=16H_in=112W_in=112kernel_size=[1,2,2]stride=[1,2,2]padding=[0,0,0]c3d-pool1
  • W2dtype=f16N=4C=128D_in=16H_in=56W_in=56kernel_size=[2,2,2]stride=[2,2,2]padding=[0,0,0]c3d-pool2
  • W3dtype=bf16N=2C=64D_in=32H_in=112W_in=112kernel_size=[3,3,3]stride=[2,2,2]padding=[1,1,1]ceil_mode=truemedicalnet-stem
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.00×240.7824torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.9711
0.7434
···
W20.00×64.8989torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
0.3008
0.1850
···
W30.00×558.8976torch_compile_ge torch_compile
torch_npu eager vendor library basis ⓘ
2.4289
1.2630
···

Positional Encoding

RopeNeoxPositionIdsFwd

x: [num_tokens, num_heads, head_dim]position_ids: [num_tokens]

num_heads=32head_dim=128

  • W1dtype=f16num_tokens=2048max_position=4096position-ids-s2k-h32-d128
  • W2dtype=bf16num_tokens=4096max_position=8192position-ids-s4k-h32-d128
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.85×0.1086ops-transformer-b2:aclnnRopeWithSinCosCache open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.2000
0.2979
0.2835
···
W21.91×0.2024ops-transformer-b2:aclnnRopeWithSinCosCache open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.3875
0.4447
0.4719
···

RopeNeoxFwd

  • W1dtype=f16seq_len=2048head_dim=64neox-1d-2k-d64
  • W2dtype=bf16seq_len=4096head_dim=128neox-1d-4k-d128
  • W3dtype=f16batch=2seq_len=2048num_heads=32head_dim=128layout=2dneox-2d-b2-s2k-h32-d128
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.14×0.0143ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.0118
0.0698
0.0638
···
W20.96×0.0173ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.0105
0.0757
0.0719
···
W30.91×0.1145ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.0949
0.3815
0.3954
···

RopeLlama31Fwd

seq_len=8192head_dim=128

  • W1dtype=bf16llama31-1d-8k-d128
  • W2dtype=f16batch=1num_heads=32layout=2dllama31-2d-b1-s8k-h32-d128
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W11.04×0.0206ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.0155
0.0921
0.0780
···
W20.86×0.2246ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library ⓘ
0.1874
1.0653
1.1926
···

RopeYarnFwd

seq_len=8192head_dim=128

  • W1dtype=bf16yarn-1d-8k-d128
  • W2dtype=f16batch=1num_heads=32layout=2dyarn-2d-b1-s8k-h32-d128
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.99×0.0209ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.0152
0.0925
0.0772
···
W20.86×0.2286ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1898
1.0504
1.1845
···

RopeLongRopeFwd

seq_len=8192head_dim=128

  • W1dtype=bf16longrope-1d-8k-d128
  • W2dtype=f16batch=1num_heads=32layout=2dlongrope-2d-b1-s8k-h32-d128
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.94×0.0203ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.0150
0.0920
0.0790
···
W20.86×0.2300ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1908
1.0408
1.2095
···

RopeNonNeoxFwd

seq_len=2048

  • W1dtype=f16head_dim=64non-neox-1d-2k-d64
  • W2dtype=bf16batch=2num_heads=32head_dim=128layout=2dnon-neox-2d-b2-s2k-h32-d128
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W10.91×0.0170ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.00813
0.0938
0.1526
···
W20.72×0.1965ops-transformer-b1:aclnnRotaryPositionEmbedding open-source ⓘ basis
torch_compile_aclgraph torch_compile
torch_npu eager vendor library
0.1358
0.4402
0.5316
···

MHC

MHCPostFwd

x_layer_out: [batch, c_x]h_post: [batch, n_expand], f32x_res: [batch, 10240]

dtype=bf16batch=4c_x=2560n_expand=4

  • W1post-large

x_layer_out: [batch, c_x]h_post: [batch, n_expand], f32x_res: [batch, 7680]

dtype=bf16batch=2c_x=1920n_expand=4

  • W2post-medium

x_layer_out: [batch, c_x]h_post: [batch, n_expand], f32x_res: [batch, 5120]

dtype=bf16batch=1c_x=1280n_expand=4

  • W3post-small
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name · tier ms TFLOP/s of ceiling by
W13.22×0.0112torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0362
0.0278
0.0296
···
W22.82×0.0085torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0240
0.0195
0.0170
···
W33.83×0.0045torch_compile_aclgraph torch_compile
torch_compile_ge torch_compile
torch_npu eager vendor library basis
0.0195
0.0123
0.0123
···

MHCPreFwd

phi: [10240, 24], f32x: [batch, 10240]b: [24], f32

dtype=bf16batch=4alpha_pre=1alpha_post=1alpha_res=1sinkhorn_repeat=1

  • W1pre-large

phi: [7680, 24], f32x: [batch, 7680]b: [24], f32

dtype=bf16batch=2alpha_pre=1alpha_post=1alpha_res=1sinkhorn_repeat=1

  • W2pre-medium

phi: [5120, 24], f32x: [batch, 5120]b: [24], f32

dtype=bf16batch=1alpha_pre=1alpha_post=1alpha_res=1sinkhorn_repeat=1

  • W3pre-small
Workload Ratio Device time Alternatives Throughput SOL Bound
alt / ours ms name ms TFLOP/s of ceiling by
W11.06×0.2026Torch-NPU eager exact reference0.2150···
W21.03×0.1253Torch-NPU eager exact reference0.1295···
W31.14×0.0757Torch-NPU eager exact reference0.0860···